Pith. sign in

Paper Citation Record · LEDGER

Listen, Think, and Understand

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 29 inbound Pith citation observations for arXiv:2305.10790.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2305.10790 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 29 of 29 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 29 of 29 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T15:38:17.057432Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

18
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation ae54b952-a671-44b9-a4dd-944c328f4b6e · inbound

SALMONN: Towards Generic Hearing Abilities for Large Language Models cites this paper.

SALMONN: Towards Generic Hearing Abilities for Large Language Models Listen, Think, and Understand

Reference 102

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T02:29:46.312500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-18T02:29:46.242983Z digest=sha256:cb6be5fac76285acb4bc49c3ab2c265faf9c6e924423a148e63e9a59fe96a4d8

Observation cd77a381-8ad5-4801-b6df-7d80726574a8 · inbound

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs cites this paper.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Listen, Think, and Understand

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.057432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.057432Z digest=sha256:cd1336e200e2a5e86719386e0154bf8e7e7bea9488daf07a2cbc3b2503a6e9d9

Observation 1bcd468c-9691-4cef-ad2d-f78d419acf2b · inbound

Audio Jailbreak: An Open Comprehensive Benchmark for Jailbreaking Large Audio-Language Models cites this paper.

Audio Jailbreak: An Open Comprehensive Benchmark for Jailbreaking Large Audio-Language Models Listen, Think, and Understand

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T15:23:01.487032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:23:01.487032Z digest=sha256:ee520139493440a4e63f6859d02bcf8eead5b6f0ba15784ce984165a9b3f8484

Observation 2dc77620-fe8f-4b0e-9d6d-95e4b0b6e9f9 · inbound

LiSTEN: Learning Soft Token Embeddings for Neural Audio LLMs cites this paper.

LiSTEN: Learning Soft Token Embeddings for Neural Audio LLMs Listen, Think, and Understand

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T14:33:46.290154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:33:46.290154Z digest=sha256:93022f73ddd24384ff7fc8961108a2d4f3584f9ad691f98f7d3c8793493418a8

Observation 4ee1ee6e-52f7-4975-adb3-778e897e1cef · inbound

ALAS: An Automatic Latent Alignment Score for Audio Language Models cites this paper.

ALAS: An Automatic Latent Alignment Score for Audio Language Models Listen, Think, and Understand

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T14:07:22.512637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:07:22.512637Z digest=sha256:951b4247c4758cd31a4d4cd99cca1b9f1d7585325c38813f6cc1346df500a81c

Observation 9f77e224-6bc4-4f47-9b94-1564d3bedd32 · inbound

ZeroSep: Separate Anything in Audio with Zero Training cites this paper.

ZeroSep: Separate Anything in Audio with Zero Training Listen, Think, and Understand

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T12:46:43.933695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:46:43.933695Z digest=sha256:2b5569d7ee0608e89c311fd5310dca25347882f021f737a951745b80eea55a7d

Observation b19e6aaa-7840-46aa-9110-6e641672ca3c · inbound

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing cites this paper.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Listen, Think, and Understand

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:52.670777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:52.670777Z digest=sha256:48f8716fd09786216b554244f2bcffcd38af0ff29df6421a674cc525f6e79f5b

Observation 552328a7-1953-4cf0-8c24-8804d68a40a1 · inbound

Teaching Physical Awareness to LLMs through Sounds cites this paper.

Teaching Physical Awareness to LLMs through Sounds Listen, Think, and Understand

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T05:15:46.367352Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:15:46.367352Z digest=sha256:4ae66f15477b256cce609bc64675d0d65d712fae4119514125a3c7df62240538

Observation ed050362-11eb-4766-9b85-376437709e6e · inbound

CoLMbo: Speaker Language Model for Descriptive Profiling cites this paper.

CoLMbo: Speaker Language Model for Descriptive Profiling Listen, Think, and Understand

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T04:54:13.382210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:54:13.382210Z digest=sha256:d306e879083db372f8ec6479144ccb57ca0b05e85462af051b572e38f538d776

Observation da6c1df5-b888-478a-a255-ece4e2b54297 · inbound

Thinking in Directivity: Speech Large Language Model for Multi-Talker Directional Speech Recognition cites this paper.

Thinking in Directivity: Speech Large Language Model for Multi-Talker Directional Speech Recognition Listen, Think, and Understand

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T00:15:33.044745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:15:33.044745Z digest=sha256:b888a0c56610c66c2e52f3fa00fef3dfe6b050b72e0d3f677d768a419bf848a2

Observation 4862910b-5f00-43ed-aa02-27eb8444c0a2 · inbound

MixAssist: An Audio-Language Dataset for Co-Creative AI Assistance in Music Mixing cites this paper.

MixAssist: An Audio-Language Dataset for Co-Creative AI Assistance in Music Mixing Listen, Think, and Understand

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T19:12:02.588908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:12:02.588908Z digest=sha256:467bd59bed91a5b8785113b2ddf889fbefacd953e71ad39488c1a674ffc64347

Observation 74ecce0c-3ad2-4bb4-8d8d-44359d549937 · inbound

Step-Audio 2 Technical Report cites this paper.

Step-Audio 2 Technical Report Listen, Think, and Understand

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-16T05:59:51.125480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T05:59:50.900436Z digest=sha256:24869f466ebfae9f9e52cfe46ca29fa7c33dcd5de81df356c2087e5b5c984292

Observation 003dd5a5-a870-433e-8397-f892f8ba0cb1 · inbound

Speech Discrete Tokens or Continuous Features? A Comparative Analysis for Spoken Language Understanding in SpeechLLMs cites this paper.

Speech Discrete Tokens or Continuous Features? A Comparative Analysis for Spoken Language Understanding in SpeechLLMs Listen, Think, and Understand

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-05T16:46:48.796047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:46:48.796047Z digest=sha256:16d4d341e80bc926bf528a2f484e03c7143ad116ba721840d5937d658c371fd0

Observation bac6bdd2-ca36-427f-a2f2-b3064ac08d97 · inbound

Improving Audio Event Recognition with Consistency Regularization cites this paper.

Improving Audio Event Recognition with Consistency Regularization Listen, Think, and Understand

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-04T17:54:57.049705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T17:54:57.049705Z digest=sha256:137222ff58b0e3c424a791c0a20823b081eaeadfc736681169f1fc1ba84ba288

Observation 589131f7-a8e1-4d51-9804-9259c77714ec · inbound

AQA-TTRL: Self-Adaptation in Audio Question Answering with Test-Time Reinforcement Learning cites this paper.

AQA-TTRL: Self-Adaptation in Audio Question Answering with Test-Time Reinforcement Learning Listen, Think, and Understand

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:17.396777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:17.396777Z digest=sha256:57787cc330fd110c04ef706cd80a27a1cb36c7ea249db3118131ad0b56e79c40

Observation 0c8682ac-6584-4734-bcc3-398ed0e6e7e3 · inbound

Audio-Cogito: Towards Deep Audio Reasoning in Large Audio Language Models cites this paper.

Audio-Cogito: Towards Deep Audio Reasoning in Large Audio Language Models Listen, Think, and Understand

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-10T14:10:28.361593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T14:10:03.707886Z digest=sha256:77e6e3767ccc076e5a2791316d9e48d9339b313b93e74181ae98d46817db5297

Observation 229861e9-348b-4d92-95c9-344776e2ebe3 · inbound

Audio-Cogito: Towards Deep Audio Reasoning in Large Audio Language Models cites this paper.

Audio-Cogito: Towards Deep Audio Reasoning in Large Audio Language Models Listen, Think, and Understand

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-12T21:18:46.566338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T21:18:46.566338Z digest=sha256:66c509a655eaa5c0f008197d04af9b5ea07da2aebc00203abca8b551b3214c93

Observation 879e6f45-04f2-405d-a31d-8cc5cf3ef6a5 · inbound

Heterogeneity-Aware Dataset Scheduling for Efficient Audio Large Language Model Training cites this paper.

Heterogeneity-Aware Dataset Scheduling for Efficient Audio Large Language Model Training Listen, Think, and Understand

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T07:18:07.291974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-20T07:13:29.041056Z digest=sha256:dbc5a79c3121da8c85d3bf8a3c4ac716f2773c7f2ae65f2fc48de0f72cb6c49f

Observation 812223e5-8250-45f0-81da-584609534252 · inbound

Codec-Robust Attacks on Audio LLMs cites this paper.

Codec-Robust Attacks on Audio LLMs Listen, Think, and Understand

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-05-21T06:44:00.649442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-21T06:43:52.735211Z digest=sha256:31e7313337ca7f31bde445c55770e8a6e245966aa33255b880f463592849069f

Observation 680c47f9-08d4-44d4-b32f-47f012750bae · inbound

Codec-Robust Attacks on Audio LLMs cites this paper.

Codec-Robust Attacks on Audio LLMs Listen, Think, and Understand

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-05-25T05:45:23.640485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-25T05:44:46.831360Z digest=sha256:68d3e96036d6c06f47524e7ecd96c77b6e885d23cc99ca36c60c521ee593ed62

Observation f870973a-d165-4d40-a668-bba16ed9cd2c · inbound

From Senses to Decisions: The Information Flow of Auditory and Visual Perception in Multimodal LLMs cites this paper.

From Senses to Decisions: The Information Flow of Auditory and Visual Perception in Multimodal LLMs Listen, Think, and Understand

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-07-03T01:57:32.387037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T16:12:53.387567Z digest=sha256:ac2dc682e18b7cd55d17610d6de3406f82f96ee4ade4d940e89b7db72a8da412

Observation 66c1f981-7254-4d5d-899a-28b74ca0c293 · inbound

Afrispeech Semantics: Evaluating Audio Semantic Reasoning in Spoken Language Models Across Domains and Accents cites this paper.

Afrispeech Semantics: Evaluating Audio Semantic Reasoning in Spoken Language Models Across Domains and Accents Listen, Think, and Understand

Reference 67

Resolution
verified exact
arxiv_id, observed 2026-06-30T22:15:05.497306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-30T22:11:44.891731Z digest=sha256:ed4db3a7b8273d97e6acbc118953f35d3e8462d5a0ec57b1dedb905211306f98

Observation a30199de-f6a5-4b06-947d-5661731f3b2b · inbound

Continuous Audio Thinking for Large Audio Language Models cites this paper.

Continuous Audio Thinking for Large Audio Language Models Listen, Think, and Understand

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T17:47:17.990037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T21:52:49.839901Z digest=sha256:2017168e327d144cc9380a84bc267ee730687eb109622c1cd92699e7d7deaeef

Observation 85c02625-1456-4d85-8dbf-f919989e25da · inbound

From Sounds to Scenes: A Benchmark for Evaluating Context-Aware Auditory Scene Understanding in Large Audio Language Models cites this paper.

From Sounds to Scenes: A Benchmark for Evaluating Context-Aware Auditory Scene Understanding in Large Audio Language Models Listen, Think, and Understand

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T20:10:07.985482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-25T20:36:24.901454Z digest=sha256:6b18f099ae26f5274f24d6d6138b8969cec7096360bebb984e2c75fab3f975bc

Observation ce9e0568-b53f-4f2a-9712-b936fe1e0f0c · inbound

Adaptive Perturbation Selection for Contrastive Audio Decoding cites this paper.

Adaptive Perturbation Selection for Contrastive Audio Decoding Listen, Think, and Understand

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:07:12.186524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-02T16:59:50.615875Z digest=sha256:e4d35deac16c56c313844ad7ed68fb516b7fdd9e7de948fe223d969df20c66d6

Observation a95e9065-c56c-4d0a-ac26-03416fe65a3e · inbound

Adaptive Perturbation Selection for Contrastive Audio Decoding cites this paper.

Adaptive Perturbation Selection for Contrastive Audio Decoding Listen, Think, and Understand

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T04:34:50.710186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:34:50.710186Z digest=sha256:f1aa5d9fd121d4e5fb103925146b0ffe3c0486904c2a8dc1d6cad4e936b604e9

Observation 7250104c-9e57-4d1b-890a-bd66ef6cea78 · inbound

Unified Audio Intelligence Without Regressing on Text Intelligence cites this paper.

Unified Audio Intelligence Without Regressing on Text Intelligence Listen, Think, and Understand

Reference 24

Resolution
metadata mismatch
local_arxiv, observed 2026-07-08T00:04:22.649038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-07-07T23:59:38.702609Z digest=sha256:7d894e4b65cc153d049ba8f75ca23c7278281c10ae2b557a30bdea30c2328b2c

Observation c1c12e22-8a62-4bc5-9aac-0135690ebf68 · inbound

Unified Audio Intelligence Without Regressing on Text Intelligence cites this paper.

Unified Audio Intelligence Without Regressing on Text Intelligence Listen, Think, and Understand

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-11T07:46:49.059192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T07:46:49.059192Z digest=sha256:1181b2213a91d531b43826801dce68c1eac3dc9613a1f9c0031f4a2115df5bf2

Observation e780d499-6b41-4c88-aca9-5acb343089f9 · inbound

Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos cites this paper.

Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos Listen, Think, and Understand

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-01T21:22:41.155465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T21:22:41.155465Z digest=sha256:34e1dda832353eac746c6502482a9c0ecc61a33a0fa08cd1ca4f358444c2c4ec