Pith. sign in

Paper Citation Record · LEDGER

GAMA: A Large Audio-Language Model with Advanced Audio Understanding and Complex Reasoning Abilities

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 24 inbound Pith citation observations for arXiv:2406.11768.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.11768 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 24 of 24 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 24 of 24 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:29:06.347770Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 6988f0d0-69cc-4727-896c-f2bdbc257da9 · inbound

Kimi-Audio Technical Report cites this paper.

Kimi-Audio Technical Report GAMA: A Large Audio-Language Model with Advanced Audio Understanding and Complex Reasoning Abilities

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:21:27.293952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T19:21:26.933349Z digest=sha256:1ec05bbd68b18828659e0878559425f99b21ca23ff3b0c31a6a64e4d3a3a2870

Observation be7bd3bd-d8fd-4713-8ca7-15a84fd921ec · inbound

Towards Spoken Mathematical Reasoning: Benchmarking Speech-based Models over Multi-faceted Math Problems cites this paper.

Towards Spoken Mathematical Reasoning: Benchmarking Speech-based Models over Multi-faceted Math Problems GAMA: A Large Audio-Language Model with Advanced Audio Understanding and Complex Reasoning Abilities

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T15:29:06.347770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:29:06.347770Z digest=sha256:bdf3466874fac048050b654818960c8ec0a537437d234c72f1831568e9494e47

Observation 7736233f-1a62-4215-91ab-9ca5f3ef4338 · inbound

Audio Jailbreak: An Open Comprehensive Benchmark for Jailbreaking Large Audio-Language Models cites this paper.

Audio Jailbreak: An Open Comprehensive Benchmark for Jailbreaking Large Audio-Language Models GAMA: A Large Audio-Language Model with Advanced Audio Understanding and Complex Reasoning Abilities

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T15:23:01.440392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:23:01.440392Z digest=sha256:6d8adedde7e32be2ba69d3e51d2893b298110abaa3e395997ff80eef7ba2411b

Observation 06d1d233-677c-4da6-a2ad-61236a8de789 · inbound

LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning cites this paper.

LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning GAMA: A Large Audio-Language Model with Advanced Audio Understanding and Complex Reasoning Abilities

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-17T03:46:06.395741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-17T03:46:06.074416Z digest=sha256:605b235fb41ed72a2364441577df66e572613a0db8aa5d80c89a2e8b95932a44

Observation c0dd767a-54c1-4b52-bf78-eb6630114878 · inbound

ZeroSep: Separate Anything in Audio with Zero Training cites this paper.

ZeroSep: Separate Anything in Audio with Zero Training GAMA: A Large Audio-Language Model with Advanced Audio Understanding and Complex Reasoning Abilities

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T12:46:44.029065Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:46:44.029065Z digest=sha256:b05d2cb55944f2ac2bd1c1e7879b5f5f2d76385dacd77322ca5e6bd8a3a32559

Observation 9c967c50-3543-44f4-b64c-13e5197827fb · inbound

CoLMbo: Speaker Language Model for Descriptive Profiling cites this paper.

CoLMbo: Speaker Language Model for Descriptive Profiling GAMA: A Large Audio-Language Model with Advanced Audio Understanding and Complex Reasoning Abilities

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T04:54:13.519625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:54:13.519625Z digest=sha256:3c3630bd2b5f93a1da9089baa13506928e3e7317bdcc0c7cef4be4d44706fe82

Observation d7d42a5d-675f-4ef8-b663-e85edb03d41f · inbound

MMW: Side Talk Rejection Multi-Microphone Whisper on Smart Glasses cites this paper.

MMW: Side Talk Rejection Multi-Microphone Whisper on Smart Glasses GAMA: A Large Audio-Language Model with Advanced Audio Understanding and Complex Reasoning Abilities

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T19:26:53.686731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:26:53.686731Z digest=sha256:b42a782daef8ac258dc65613110e1a0280c5669380db0a35f3bca4705d543535

Observation 2e732449-e696-4d5f-ae28-3f39fa1338d7 · inbound

Attacker's Noise Can Manipulate Your Audio-based LLM in the Real World cites this paper.

Attacker's Noise Can Manipulate Your Audio-based LLM in the Real World GAMA: A Large Audio-Language Model with Advanced Audio Understanding and Complex Reasoning Abilities

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T19:46:11.380266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:46:11.380266Z digest=sha256:e14888111f8711d8963f4af2ceac8a74813810dc024bcda6e1e54ba2a647ea9b

Observation a4ccc679-865f-47a5-9b35-54313167a882 · inbound

The Man Behind the Sound: Demystifying Audio Private Attribute Profiling via Multimodal Large Language Model Agents cites this paper.

The Man Behind the Sound: Demystifying Audio Private Attribute Profiling via Multimodal Large Language Model Agents GAMA: A Large Audio-Language Model with Advanced Audio Understanding and Complex Reasoning Abilities

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T17:45:45.070152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:45:45.070152Z digest=sha256:3970ae576e6645e362d1334942f84daeb873c2bc030681fd46b393250bfe1f79

Observation 97846861-2b18-4eba-a1bd-0d568262121a · inbound

BoSS: Beyond-Semantic Speech cites this paper.

BoSS: Beyond-Semantic Speech GAMA: A Large Audio-Language Model with Advanced Audio Understanding and Complex Reasoning Abilities

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T14:50:04.254194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:50:04.254194Z digest=sha256:3d44e7b8f4598fb3bdc769256a23a07069e2565f7ca247ae1d6a14a936c8d8e8

Observation b03fa92e-d669-4b4c-b3cc-618d47076bdd · inbound

From Sound to Sight: Towards AI-authored Music Videos cites this paper.

From Sound to Sight: Towards AI-authored Music Videos GAMA: A Large Audio-Language Model with Advanced Audio Understanding and Complex Reasoning Abilities

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T18:23:28.522442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T18:23:28.522442Z digest=sha256:d8d290e8424cf9b5a7229a039c57f818363b940944c788eccd1d9c882683c4d6

Observation aaef49ed-b065-46dc-bb5a-1772b68ba670 · inbound

Improving Audio Event Recognition with Consistency Regularization cites this paper.

Improving Audio Event Recognition with Consistency Regularization GAMA: A Large Audio-Language Model with Advanced Audio Understanding and Complex Reasoning Abilities

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T17:54:57.124210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T17:54:57.124210Z digest=sha256:3b1ecf889141c2dc7ebbbc221355a22cf54dca25854fff0d1a32378bc438c143

Observation 9ac6ba19-b231-44c4-a136-9a515157f6f3 · inbound

See, Hear, and Understand: Benchmarking Audiovisual Human Speech Understanding in Multimodal Large Language Models cites this paper.

See, Hear, and Understand: Benchmarking Audiovisual Human Speech Understanding in Multimodal Large Language Models GAMA: A Large Audio-Language Model with Advanced Audio Understanding and Complex Reasoning Abilities

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:13:52.187373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-17T02:12:55.170296Z digest=sha256:71b8da3b1b30c04184509dfd3411bb52af607a3b7da039a8b5878a9940e651e8

Observation b8a97e48-dde2-4ad5-8b68-ff191afc5bd4 · inbound

FastSLM: Hierarchical Temporal Abstraction for Efficient Long-Form Speech Adaptation cites this paper.

FastSLM: Hierarchical Temporal Abstraction for Efficient Long-Form Speech Adaptation GAMA: A Large Audio-Language Model with Advanced Audio Understanding and Complex Reasoning Abilities

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-03T12:01:59.901861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:01:59.901861Z digest=sha256:c010fc41c84e5ab3f17192d1b1b1596be47680035c0a70aa82d7aef6cc15352e

Observation 807caf8f-a394-43e0-8e96-8cd188daff0d · inbound

EvA: An Evidence-First Audio Understanding Paradigm for LALMs cites this paper.

EvA: An Evidence-First Audio Understanding Paradigm for LALMs GAMA: A Large Audio-Language Model with Advanced Audio Understanding and Complex Reasoning Abilities

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-13T16:50:28.996341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T16:50:28.996341Z digest=sha256:56b18d7b769a25f904cfb644c55fe9fae3c000541c17ccb2a8568060ed9958b7

Observation 2897d192-ccdc-40cd-ae15-6f9526583675 · inbound

Jamendo-MT-QA: A Benchmark for Multi-Track Comparative Music Question Answering cites this paper.

Jamendo-MT-QA: A Benchmark for Multi-Track Comparative Music Question Answering GAMA: A Large Audio-Language Model with Advanced Audio Understanding and Complex Reasoning Abilities

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T07:10:58.086055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T17:16:59.034480Z digest=sha256:aabad71ac7dc07d3f01a645fc92cc5b6ca2b96f9d3dc4b472bdecb73a9b76588

Observation e8da2d69-0366-4eef-81b8-3841b3aeff5e · inbound

Listen, Pause, and Reason: Toward Perception-Grounded Hybrid Reasoning for Audio Understanding cites this paper.

Listen, Pause, and Reason: Toward Perception-Grounded Hybrid Reasoning for Audio Understanding GAMA: A Large Audio-Language Model with Advanced Audio Understanding and Complex Reasoning Abilities

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T09:28:38.897129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T09:24:25.616750Z digest=sha256:575f119fdf8d6abb279624bef70f85d3ecdcfa2a9f2d7dbd4f28b88d367698fe

Observation d74d148d-6114-4231-a16b-f7a6508132cf · inbound

Temporal Contrastive Decoding: A Training-Free Method for Large Audio-Language Models cites this paper.

Temporal Contrastive Decoding: A Training-Free Method for Large Audio-Language Models GAMA: A Large Audio-Language Model with Advanced Audio Understanding and Complex Reasoning Abilities

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T10:19:20.685275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T10:16:26.526986Z digest=sha256:3e652500b488a82cb3d39bbed4e139714d1790486ab2e90d6fbc60a0e91c1bcc

Observation 34b70cd5-64b5-4786-8ee0-e3be2ff31c97 · inbound

Frame-Aligned Fusion of Canary and WavLM for Non-Intrusive Intelligibility Prediction of Hearing-Aid-Processed Speech cites this paper.

Frame-Aligned Fusion of Canary and WavLM for Non-Intrusive Intelligibility Prediction of Hearing-Aid-Processed Speech GAMA: A Large Audio-Language Model with Advanced Audio Understanding and Complex Reasoning Abilities

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-25T02:35:14.575239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-25T02:34:11.297579Z digest=sha256:182f77dca361a058c28f6abb56798e5384898deda8421b6a3de3f925ea3fdcba

Observation 6fed932d-bdcb-4364-b48d-b905fc9f6aa3 · inbound

MCBench: A Multicontext Safety Assessment Benchmark for Omni Large Language Models cites this paper.

MCBench: A Multicontext Safety Assessment Benchmark for Omni Large Language Models GAMA: A Large Audio-Language Model with Advanced Audio Understanding and Complex Reasoning Abilities

Reference 41

Resolution
unresolved
no resolver link, observed 2026-07-12T19:28:32.792370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T19:28:32.792370Z digest=sha256:07a42a9a68bea217267dcd1b547284334b7ac8b3f2974bd75892ab1e6b29032d

Observation b41efbd5-ea20-4409-a767-f73043369a62 · inbound

TinyGiantALM: A Compact Audio-Language Model for Intent-Aware Reasoning under Resource Constraints cites this paper.

TinyGiantALM: A Compact Audio-Language Model for Intent-Aware Reasoning under Resource Constraints GAMA: A Large Audio-Language Model with Advanced Audio Understanding and Complex Reasoning Abilities

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-07-02T23:17:29.830453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T18:18:41.816342Z digest=sha256:934716656b5d0d3ff94f6cc35439495d91e432a5d8e43b5ec5bd6629c40ae8c7

Observation 442e0517-a4c0-4866-b694-c9b6a67c0ec6 · inbound

Afrispeech Semantics: Evaluating Audio Semantic Reasoning in Spoken Language Models Across Domains and Accents cites this paper.

Afrispeech Semantics: Evaluating Audio Semantic Reasoning in Spoken Language Models Across Domains and Accents GAMA: A Large Audio-Language Model with Advanced Audio Understanding and Complex Reasoning Abilities

Reference 70

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T22:15:05.181500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-30T22:11:44.891731Z digest=sha256:2f67d3aa2e8929e6e03f058e975fd0a0287aa012fb319389ff8f2c2ac55be05a

Observation 4f5c37bc-1184-4884-bf90-ea0a7a740cfe · inbound

Unlocking In-Context Learning in Audio-Language Models from Decentralized Medical Audio cites this paper.

Unlocking In-Context Learning in Audio-Language Models from Decentralized Medical Audio GAMA: A Large Audio-Language Model with Advanced Audio Understanding and Complex Reasoning Abilities

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-07-04T09:59:45.495210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T09:16:25.450336Z digest=sha256:d7e338b5bbf7927719b3b7536cd14ee18f9dec2d05ce7dc995580dd487c4baff

Observation c8e42549-7300-4a39-b422-21309c1f4c80 · inbound

Empowering Long-form Omni-modal Understanding with Robust Audio Perception cites this paper.

Empowering Long-form Omni-modal Understanding with Robust Audio Perception GAMA: A Large Audio-Language Model with Advanced Audio Understanding and Complex Reasoning Abilities

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-14T12:48:58.688011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T12:48:58.688011Z digest=sha256:8879d13ec95777e9bc6ca4c74eb75df51cae34be9c218a99fa75d2c80e264502