Pith. sign in

Paper Citation Record · LEDGER

GAMA: A Large Audio-Language Model with Advanced Audio Understanding and Complex Reasoning Abilities

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 27 inbound Pith citation observations for arXiv:2406.11768.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.11768 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 27 of 27 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T22:40:25.950869Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation c0617f60-88eb-4658-89b5-6cd0b76d2cf5 · inbound

ADIFF: Explaining audio difference using natural language cites this paper.

ADIFF: Explaining audio difference using natural language GAMA: A Large Audio-Language Model with Advanced Audio Understanding and Complex Reasoning Abilities

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-08T22:40:25.950869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:40:25.950869Z digest=sha256:74221d4c1bc731529074b269184605f5f010d1bfbc76d9b47518358dabab7787

Observation ba213a9a-be8e-45cf-893e-b3cbf58b5c4c · inbound

Time2Lang: Bridging Time-Series Foundation Models and Large Language Models for Health Sensing Beyond Prompting cites this paper.

Time2Lang: Bridging Time-Series Foundation Models and Large Language Models for Health Sensing Beyond Prompting GAMA: A Large Audio-Language Model with Advanced Audio Understanding and Complex Reasoning Abilities

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-08T12:14:20.285144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T12:14:20.285144Z digest=sha256:9c7fcd7946374bba6bd189abe9b711962b15a4d8fbb96bbcfeafe22c4464fb1d

Observation b145106a-5068-410f-925b-b26c182da3e3 · inbound

From No to Know: Taxonomy, Challenges, and Opportunities for Negation Understanding in Multimodal Foundation Models cites this paper.

From No to Know: Taxonomy, Challenges, and Opportunities for Negation Understanding in Multimodal Foundation Models GAMA: A Large Audio-Language Model with Advanced Audio Understanding and Complex Reasoning Abilities

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-08T14:46:05.831117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:46:05.831117Z digest=sha256:39d3dfdd267e6562795a56c6b90919a4526b955b6189b87e5e985895f85afb61

Observation 6988f0d0-69cc-4727-896c-f2bdbc257da9 · inbound

Kimi-Audio Technical Report cites this paper.

Kimi-Audio Technical Report GAMA: A Large Audio-Language Model with Advanced Audio Understanding and Complex Reasoning Abilities

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:21:27.293952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-11T19:21:26.933349Z digest=sha256:dd70033256e60637d11b2eefe00a250484de3c3e4de92711ad7049a80cc29d40

Observation be7bd3bd-d8fd-4713-8ca7-15a84fd921ec · inbound

Towards Spoken Mathematical Reasoning: Benchmarking Speech-based Models over Multi-faceted Math Problems cites this paper.

Towards Spoken Mathematical Reasoning: Benchmarking Speech-based Models over Multi-faceted Math Problems GAMA: A Large Audio-Language Model with Advanced Audio Understanding and Complex Reasoning Abilities

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T15:29:06.347770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:29:06.347770Z digest=sha256:bdf3466874fac048050b654818960c8ec0a537437d234c72f1831568e9494e47

Observation 7736233f-1a62-4215-91ab-9ca5f3ef4338 · inbound

Audio Jailbreak: An Open Comprehensive Benchmark for Jailbreaking Large Audio-Language Models cites this paper.

Audio Jailbreak: An Open Comprehensive Benchmark for Jailbreaking Large Audio-Language Models GAMA: A Large Audio-Language Model with Advanced Audio Understanding and Complex Reasoning Abilities

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T15:23:01.440392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:23:01.440392Z digest=sha256:f0ef4096f6958a82430dc7f4cb2266efde6e097032ee6e17fa02b1c06426b1c3

Observation 06d1d233-677c-4da6-a2ad-61236a8de789 · inbound

LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning cites this paper.

LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning GAMA: A Large Audio-Language Model with Advanced Audio Understanding and Complex Reasoning Abilities

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-17T03:46:06.395741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T03:46:06.074416Z digest=sha256:b076d240d0683c468625922c139a05b60e267c1d5013fc46388da9d6cefb9e03

Observation c0dd767a-54c1-4b52-bf78-eb6630114878 · inbound

ZeroSep: Separate Anything in Audio with Zero Training cites this paper.

ZeroSep: Separate Anything in Audio with Zero Training GAMA: A Large Audio-Language Model with Advanced Audio Understanding and Complex Reasoning Abilities

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T12:46:44.029065Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:46:44.029065Z digest=sha256:db2a63594a415217b4947d2e66d60b91dce25310fa1feebefd531dc16748eb8a

Observation 9c967c50-3543-44f4-b64c-13e5197827fb · inbound

CoLMbo: Speaker Language Model for Descriptive Profiling cites this paper.

CoLMbo: Speaker Language Model for Descriptive Profiling GAMA: A Large Audio-Language Model with Advanced Audio Understanding and Complex Reasoning Abilities

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T04:54:13.519625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:54:13.519625Z digest=sha256:3c3630bd2b5f93a1da9089baa13506928e3e7317bdcc0c7cef4be4d44706fe82

Observation d7d42a5d-675f-4ef8-b663-e85edb03d41f · inbound

MMW: Side Talk Rejection Multi-Microphone Whisper on Smart Glasses cites this paper.

MMW: Side Talk Rejection Multi-Microphone Whisper on Smart Glasses GAMA: A Large Audio-Language Model with Advanced Audio Understanding and Complex Reasoning Abilities

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T19:26:53.686731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:26:53.686731Z digest=sha256:6f99bd48554750d212432f07314836f2ccbb01c1c71c553abc9459e5c2c94b4b

Observation 2e732449-e696-4d5f-ae28-3f39fa1338d7 · inbound

Attacker's Noise Can Manipulate Your Audio-based LLM in the Real World cites this paper.

Attacker's Noise Can Manipulate Your Audio-based LLM in the Real World GAMA: A Large Audio-Language Model with Advanced Audio Understanding and Complex Reasoning Abilities

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T19:46:11.380266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:46:11.380266Z digest=sha256:e9e70b85bbb577aa6fdcb6a643ab6e9d921d8bb69fd30471d632658aa610b20a

Observation a4ccc679-865f-47a5-9b35-54313167a882 · inbound

The Man Behind the Sound: Demystifying Audio Private Attribute Profiling via Multimodal Large Language Model Agents cites this paper.

The Man Behind the Sound: Demystifying Audio Private Attribute Profiling via Multimodal Large Language Model Agents GAMA: A Large Audio-Language Model with Advanced Audio Understanding and Complex Reasoning Abilities

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T17:45:45.070152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:45:45.070152Z digest=sha256:8546339d016716eefb24b36fd8ac9719b29a42d1f05501329887d91076c139f6

Observation 97846861-2b18-4eba-a1bd-0d568262121a · inbound

BoSS: Beyond-Semantic Speech cites this paper.

BoSS: Beyond-Semantic Speech GAMA: A Large Audio-Language Model with Advanced Audio Understanding and Complex Reasoning Abilities

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T14:50:04.254194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:50:04.254194Z digest=sha256:7fb4a5709fe487f6a6f7ad8a629f5035c7d6363569edae28cbb2deaed54f6760

Observation b03fa92e-d669-4b4c-b3cc-618d47076bdd · inbound

From Sound to Sight: Towards AI-authored Music Videos cites this paper.

From Sound to Sight: Towards AI-authored Music Videos GAMA: A Large Audio-Language Model with Advanced Audio Understanding and Complex Reasoning Abilities

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T18:23:28.522442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T18:23:28.522442Z digest=sha256:f79668e3ce1b39af49015366d2cf1c593fccbbe714cfdc01f9b03a44b844337c

Observation aaef49ed-b065-46dc-bb5a-1772b68ba670 · inbound

Improving Audio Event Recognition with Consistency Regularization cites this paper.

Improving Audio Event Recognition with Consistency Regularization GAMA: A Large Audio-Language Model with Advanced Audio Understanding and Complex Reasoning Abilities

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T17:54:57.124210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T17:54:57.124210Z digest=sha256:3b1ecf889141c2dc7ebbbc221355a22cf54dca25854fff0d1a32378bc438c143

Observation 9ac6ba19-b231-44c4-a136-9a515157f6f3 · inbound

See, Hear, and Understand: Benchmarking Audiovisual Human Speech Understanding in Multimodal Large Language Models cites this paper.

See, Hear, and Understand: Benchmarking Audiovisual Human Speech Understanding in Multimodal Large Language Models GAMA: A Large Audio-Language Model with Advanced Audio Understanding and Complex Reasoning Abilities

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:13:52.187373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T02:12:55.170296Z digest=sha256:5b297a066eaf0a5ee363631ddc6bd2bc0135241df6a011b64615d7dddd831b63

Observation b8a97e48-dde2-4ad5-8b68-ff191afc5bd4 · inbound

FastSLM: Hierarchical Temporal Abstraction for Efficient Long-Form Speech Adaptation cites this paper.

FastSLM: Hierarchical Temporal Abstraction for Efficient Long-Form Speech Adaptation GAMA: A Large Audio-Language Model with Advanced Audio Understanding and Complex Reasoning Abilities

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-03T12:01:59.901861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:01:59.901861Z digest=sha256:c010fc41c84e5ab3f17192d1b1b1596be47680035c0a70aa82d7aef6cc15352e

Observation 807caf8f-a394-43e0-8e96-8cd188daff0d · inbound

EvA: An Evidence-First Audio Understanding Paradigm for LALMs cites this paper.

EvA: An Evidence-First Audio Understanding Paradigm for LALMs GAMA: A Large Audio-Language Model with Advanced Audio Understanding and Complex Reasoning Abilities

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-13T16:50:28.996341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T16:50:28.996341Z digest=sha256:56b18d7b769a25f904cfb644c55fe9fae3c000541c17ccb2a8568060ed9958b7

Observation 2897d192-ccdc-40cd-ae15-6f9526583675 · inbound

Jamendo-MT-QA: A Benchmark for Multi-Track Comparative Music Question Answering cites this paper.

Jamendo-MT-QA: A Benchmark for Multi-Track Comparative Music Question Answering GAMA: A Large Audio-Language Model with Advanced Audio Understanding and Complex Reasoning Abilities

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T07:10:58.086055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T17:16:59.034480Z digest=sha256:529f2b34648ea0ab06323d3030a3e8db692254b20ba96ade6576596145c700f7

Observation e8da2d69-0366-4eef-81b8-3841b3aeff5e · inbound

Listen, Pause, and Reason: Toward Perception-Grounded Hybrid Reasoning for Audio Understanding cites this paper.

Listen, Pause, and Reason: Toward Perception-Grounded Hybrid Reasoning for Audio Understanding GAMA: A Large Audio-Language Model with Advanced Audio Understanding and Complex Reasoning Abilities

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T09:28:38.897129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T09:24:25.616750Z digest=sha256:ece58c603c157c54c65d08fc70e031b8a633d12144919126c58c6668d9ec37d8

Observation d74d148d-6114-4231-a16b-f7a6508132cf · inbound

Temporal Contrastive Decoding: A Training-Free Method for Large Audio-Language Models cites this paper.

Temporal Contrastive Decoding: A Training-Free Method for Large Audio-Language Models GAMA: A Large Audio-Language Model with Advanced Audio Understanding and Complex Reasoning Abilities

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T10:19:20.685275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T10:16:26.526986Z digest=sha256:3af3d3dabc22cba4c2bf8ab9718599128d16d175726843f75998d66ca96785f2

Observation 34b70cd5-64b5-4786-8ee0-e3be2ff31c97 · inbound

Frame-Aligned Fusion of Canary and WavLM for Non-Intrusive Intelligibility Prediction of Hearing-Aid-Processed Speech cites this paper.

Frame-Aligned Fusion of Canary and WavLM for Non-Intrusive Intelligibility Prediction of Hearing-Aid-Processed Speech GAMA: A Large Audio-Language Model with Advanced Audio Understanding and Complex Reasoning Abilities

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-25T02:35:14.575239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T02:34:11.297579Z digest=sha256:e45702bd0f5d7c1058a63cfc81a3cf6d3600c9ec74bad10b5f4067c98637c031

Observation 6fed932d-bdcb-4364-b48d-b905fc9f6aa3 · inbound

MCBench: A Multicontext Safety Assessment Benchmark for Omni Large Language Models cites this paper.

MCBench: A Multicontext Safety Assessment Benchmark for Omni Large Language Models GAMA: A Large Audio-Language Model with Advanced Audio Understanding and Complex Reasoning Abilities

Reference 41

Resolution
unresolved
no resolver link, observed 2026-07-12T19:28:32.792370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T19:28:32.792370Z digest=sha256:07a42a9a68bea217267dcd1b547284334b7ac8b3f2974bd75892ab1e6b29032d

Observation b41efbd5-ea20-4409-a767-f73043369a62 · inbound

TinyGiantALM: A Compact Audio-Language Model for Intent-Aware Reasoning under Resource Constraints cites this paper.

TinyGiantALM: A Compact Audio-Language Model for Intent-Aware Reasoning under Resource Constraints GAMA: A Large Audio-Language Model with Advanced Audio Understanding and Complex Reasoning Abilities

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-07-02T23:17:29.830453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T18:18:41.816342Z digest=sha256:6b8e7dac0e6f53e4cc9ff0a9559fc6d7b8b0dac3a4b96a2195a610ad3ab13a11

Observation 442e0517-a4c0-4866-b694-c9b6a67c0ec6 · inbound

Afrispeech Semantics: Evaluating Audio Semantic Reasoning in Spoken Language Models Across Domains and Accents cites this paper.

Afrispeech Semantics: Evaluating Audio Semantic Reasoning in Spoken Language Models Across Domains and Accents GAMA: A Large Audio-Language Model with Advanced Audio Understanding and Complex Reasoning Abilities

Reference 70

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T22:15:05.181500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-30T22:11:44.891731Z digest=sha256:2cb2023f577f564060ce4c6b297ff7bf7500a3777690e39de87b2da22a545b7b

Observation 4f5c37bc-1184-4884-bf90-ea0a7a740cfe · inbound

Unlocking In-Context Learning in Audio-Language Models from Decentralized Medical Audio cites this paper.

Unlocking In-Context Learning in Audio-Language Models from Decentralized Medical Audio GAMA: A Large Audio-Language Model with Advanced Audio Understanding and Complex Reasoning Abilities

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-07-04T09:59:45.495210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T09:16:25.450336Z digest=sha256:69e4ca5eded99f39aedce253eacba54c3e4bc1b48820a6ce652793fad23b3f78

Observation c8e42549-7300-4a39-b422-21309c1f4c80 · inbound

Empowering Long-form Omni-modal Understanding with Robust Audio Perception cites this paper.

Empowering Long-form Omni-modal Understanding with Robust Audio Perception GAMA: A Large Audio-Language Model with Advanced Audio Understanding and Complex Reasoning Abilities

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-14T12:48:58.688011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T12:48:58.688011Z digest=sha256:57dc01ad62241d4fa1f0c03e8082da3698d4ccdfb355f06eacb88c23612a9773