Pith. sign in

Paper Citation Record · LEDGER

Listen, Think, and Understand

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 29 inbound Pith citation observations for arXiv:2305.10790.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2305.10790 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 29 of 29 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 29 of 29 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T15:38:17.057432Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

18
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation ae54b952-a671-44b9-a4dd-944c328f4b6e · inbound

SALMONN: Towards Generic Hearing Abilities for Large Language Models cites this paper.

SALMONN: Towards Generic Hearing Abilities for Large Language Models Listen, Think, and Understand

Reference 102

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T02:29:46.312500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-18T02:29:46.242983Z digest=sha256:1528f30c5b5d71e01bf9bd0e3acc22cdfe931a77067e6a0ab22e17cc38fc5d99

Observation cd77a381-8ad5-4801-b6df-7d80726574a8 · inbound

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs cites this paper.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Listen, Think, and Understand

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.057432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.057432Z digest=sha256:f5992bb4894027399ac59fca35a5671d467b67703590307867ac4d107a3d4e89

Observation 1bcd468c-9691-4cef-ad2d-f78d419acf2b · inbound

Audio Jailbreak: An Open Comprehensive Benchmark for Jailbreaking Large Audio-Language Models cites this paper.

Audio Jailbreak: An Open Comprehensive Benchmark for Jailbreaking Large Audio-Language Models Listen, Think, and Understand

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T15:23:01.487032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:23:01.487032Z digest=sha256:3b2364dadd29511dbb51f0e1b253bf928345afb41aaec694a3d232d9ce099ae0

Observation 2dc77620-fe8f-4b0e-9d6d-95e4b0b6e9f9 · inbound

LiSTEN: Learning Soft Token Embeddings for Neural Audio LLMs cites this paper.

LiSTEN: Learning Soft Token Embeddings for Neural Audio LLMs Listen, Think, and Understand

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T14:33:46.290154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:33:46.290154Z digest=sha256:9c312f182caf47de781df9d3ff2e4bc0cef5a603dd29e7dc9625e7369ce42edd

Observation 4ee1ee6e-52f7-4975-adb3-778e897e1cef · inbound

ALAS: An Automatic Latent Alignment Score for Audio Language Models cites this paper.

ALAS: An Automatic Latent Alignment Score for Audio Language Models Listen, Think, and Understand

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T14:07:22.512637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:07:22.512637Z digest=sha256:448ca9a364354cdddbc3bb8429fcba85cf62ee0a37a573bc93c894d6a64b610b

Observation 9f77e224-6bc4-4f47-9b94-1564d3bedd32 · inbound

ZeroSep: Separate Anything in Audio with Zero Training cites this paper.

ZeroSep: Separate Anything in Audio with Zero Training Listen, Think, and Understand

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T12:46:43.933695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:46:43.933695Z digest=sha256:405da289bf16a6638f21d5d1de8f3e35a70243a80fa0cea0eef17e4f1ee761a4

Observation b19e6aaa-7840-46aa-9110-6e641672ca3c · inbound

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing cites this paper.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Listen, Think, and Understand

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:52.670777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:52.670777Z digest=sha256:aa6d0d4c9815929a23c6a22e4b208f5b317a79c153cea5f64491271129453146

Observation 552328a7-1953-4cf0-8c24-8804d68a40a1 · inbound

Teaching Physical Awareness to LLMs through Sounds cites this paper.

Teaching Physical Awareness to LLMs through Sounds Listen, Think, and Understand

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T05:15:46.367352Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:15:46.367352Z digest=sha256:4ae66f15477b256cce609bc64675d0d65d712fae4119514125a3c7df62240538

Observation ed050362-11eb-4766-9b85-376437709e6e · inbound

CoLMbo: Speaker Language Model for Descriptive Profiling cites this paper.

CoLMbo: Speaker Language Model for Descriptive Profiling Listen, Think, and Understand

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T04:54:13.382210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:54:13.382210Z digest=sha256:d306e879083db372f8ec6479144ccb57ca0b05e85462af051b572e38f538d776

Observation da6c1df5-b888-478a-a255-ece4e2b54297 · inbound

Thinking in Directivity: Speech Large Language Model for Multi-Talker Directional Speech Recognition cites this paper.

Thinking in Directivity: Speech Large Language Model for Multi-Talker Directional Speech Recognition Listen, Think, and Understand

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T00:15:33.044745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:15:33.044745Z digest=sha256:6cdc633e4b6763d077365f432169a85f13d06ccb3bd194109520f42eaae915f9

Observation 4862910b-5f00-43ed-aa02-27eb8444c0a2 · inbound

MixAssist: An Audio-Language Dataset for Co-Creative AI Assistance in Music Mixing cites this paper.

MixAssist: An Audio-Language Dataset for Co-Creative AI Assistance in Music Mixing Listen, Think, and Understand

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T19:12:02.588908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:12:02.588908Z digest=sha256:c142f6efc0b05d5feef6181ef2b599937ed13bd5c495fc466aee15a7ac179b6a

Observation 74ecce0c-3ad2-4bb4-8d8d-44359d549937 · inbound

Step-Audio 2 Technical Report cites this paper.

Step-Audio 2 Technical Report Listen, Think, and Understand

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-16T05:59:51.125480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T05:59:50.900436Z digest=sha256:cf06545eb5d1124c9128a1c83239f16c247fa584bbcae6443349607283947d49

Observation 003dd5a5-a870-433e-8397-f892f8ba0cb1 · inbound

Speech Discrete Tokens or Continuous Features? A Comparative Analysis for Spoken Language Understanding in SpeechLLMs cites this paper.

Speech Discrete Tokens or Continuous Features? A Comparative Analysis for Spoken Language Understanding in SpeechLLMs Listen, Think, and Understand

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-05T16:46:48.796047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:46:48.796047Z digest=sha256:d5409d1f74f40c5f33a2e4a40d8247eb9bd1026f208f062c93c45d23e3c02783

Observation bac6bdd2-ca36-427f-a2f2-b3064ac08d97 · inbound

Improving Audio Event Recognition with Consistency Regularization cites this paper.

Improving Audio Event Recognition with Consistency Regularization Listen, Think, and Understand

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-04T17:54:57.049705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T17:54:57.049705Z digest=sha256:7d37b0fcb68409aae7e47d8c2b415180968d51eba4cb493b640c944b06edd05c

Observation 589131f7-a8e1-4d51-9804-9259c77714ec · inbound

AQA-TTRL: Self-Adaptation in Audio Question Answering with Test-Time Reinforcement Learning cites this paper.

AQA-TTRL: Self-Adaptation in Audio Question Answering with Test-Time Reinforcement Learning Listen, Think, and Understand

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:17.396777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:17.396777Z digest=sha256:57787cc330fd110c04ef706cd80a27a1cb36c7ea249db3118131ad0b56e79c40

Observation 0c8682ac-6584-4734-bcc3-398ed0e6e7e3 · inbound

Audio-Cogito: Towards Deep Audio Reasoning in Large Audio Language Models cites this paper.

Audio-Cogito: Towards Deep Audio Reasoning in Large Audio Language Models Listen, Think, and Understand

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-10T14:10:28.361593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T14:10:03.707886Z digest=sha256:edc116678ebaaf71891abad24773f2cb01b4542193ef4cff2fd76e5f190a25e2

Observation 229861e9-348b-4d92-95c9-344776e2ebe3 · inbound

Audio-Cogito: Towards Deep Audio Reasoning in Large Audio Language Models cites this paper.

Audio-Cogito: Towards Deep Audio Reasoning in Large Audio Language Models Listen, Think, and Understand

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-12T21:18:46.566338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T21:18:46.566338Z digest=sha256:66c509a655eaa5c0f008197d04af9b5ea07da2aebc00203abca8b551b3214c93

Observation 879e6f45-04f2-405d-a31d-8cc5cf3ef6a5 · inbound

Heterogeneity-Aware Dataset Scheduling for Efficient Audio Large Language Model Training cites this paper.

Heterogeneity-Aware Dataset Scheduling for Efficient Audio Large Language Model Training Listen, Think, and Understand

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T07:18:07.291974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-20T07:13:29.041056Z digest=sha256:42a446aa42e230a57e897c7c5b6e58a7e7db700ea576a5752d92971a9c54131b

Observation 812223e5-8250-45f0-81da-584609534252 · inbound

Codec-Robust Attacks on Audio LLMs cites this paper.

Codec-Robust Attacks on Audio LLMs Listen, Think, and Understand

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-05-21T06:44:00.649442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T06:43:52.735211Z digest=sha256:e72e902fd2e283820ff0ba7af7c1821cc98e8c2580bc8c7c6516926db7704e0f

Observation 680c47f9-08d4-44d4-b32f-47f012750bae · inbound

Codec-Robust Attacks on Audio LLMs cites this paper.

Codec-Robust Attacks on Audio LLMs Listen, Think, and Understand

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-05-25T05:45:23.640485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-25T05:44:46.831360Z digest=sha256:829766ffc67918cb2c2f33f61700a84f78d54de3ecf482678c710d2504c0d2d2

Observation f870973a-d165-4d40-a668-bba16ed9cd2c · inbound

From Senses to Decisions: The Information Flow of Auditory and Visual Perception in Multimodal LLMs cites this paper.

From Senses to Decisions: The Information Flow of Auditory and Visual Perception in Multimodal LLMs Listen, Think, and Understand

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-07-03T01:57:32.387037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-27T16:12:53.387567Z digest=sha256:a17a436a9f1b0e4767d453f52b04fbc6d62fc2e3d961c5330d08cfdf3fe19816

Observation 66c1f981-7254-4d5d-899a-28b74ca0c293 · inbound

Afrispeech Semantics: Evaluating Audio Semantic Reasoning in Spoken Language Models Across Domains and Accents cites this paper.

Afrispeech Semantics: Evaluating Audio Semantic Reasoning in Spoken Language Models Across Domains and Accents Listen, Think, and Understand

Reference 67

Resolution
verified exact
arxiv_id, observed 2026-06-30T22:15:05.497306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-30T22:11:44.891731Z digest=sha256:0865baad93ed876fde36f7e356b52ba042c3b23b7f0078d13b76808d2d505ea9

Observation a30199de-f6a5-4b06-947d-5661731f3b2b · inbound

Continuous Audio Thinking for Large Audio Language Models cites this paper.

Continuous Audio Thinking for Large Audio Language Models Listen, Think, and Understand

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T17:47:17.990037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-27T21:52:49.839901Z digest=sha256:52cb184c84e742ebaf9876a6a92241ab2e58bed083929b956d10ee39a2aa89be

Observation 85c02625-1456-4d85-8dbf-f919989e25da · inbound

From Sounds to Scenes: A Benchmark for Evaluating Context-Aware Auditory Scene Understanding in Large Audio Language Models cites this paper.

From Sounds to Scenes: A Benchmark for Evaluating Context-Aware Auditory Scene Understanding in Large Audio Language Models Listen, Think, and Understand

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T20:10:07.985482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-25T20:36:24.901454Z digest=sha256:851a259141652408076717f9dd54cbee5c6755b36d92b61d6af02ab956ef7c6b

Observation ce9e0568-b53f-4f2a-9712-b936fe1e0f0c · inbound

Adaptive Perturbation Selection for Contrastive Audio Decoding cites this paper.

Adaptive Perturbation Selection for Contrastive Audio Decoding Listen, Think, and Understand

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:07:12.186524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-02T16:59:50.615875Z digest=sha256:c20bf4480811af5ee780f5c64fa751a07aafd41dd4fd282d0e02c198aed53268

Observation a95e9065-c56c-4d0a-ac26-03416fe65a3e · inbound

Adaptive Perturbation Selection for Contrastive Audio Decoding cites this paper.

Adaptive Perturbation Selection for Contrastive Audio Decoding Listen, Think, and Understand

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T04:34:50.710186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:34:50.710186Z digest=sha256:f1aa5d9fd121d4e5fb103925146b0ffe3c0486904c2a8dc1d6cad4e936b604e9

Observation 7250104c-9e57-4d1b-890a-bd66ef6cea78 · inbound

Unified Audio Intelligence Without Regressing on Text Intelligence cites this paper.

Unified Audio Intelligence Without Regressing on Text Intelligence Listen, Think, and Understand

Reference 24

Resolution
metadata mismatch
local_arxiv, observed 2026-07-08T00:04:22.649038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-07-07T23:59:38.702609Z digest=sha256:f694838c03be7af2f171007818cf247e1566161f92586d3942c68dc5837e5714

Observation c1c12e22-8a62-4bc5-9aac-0135690ebf68 · inbound

Unified Audio Intelligence Without Regressing on Text Intelligence cites this paper.

Unified Audio Intelligence Without Regressing on Text Intelligence Listen, Think, and Understand

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-11T07:46:49.059192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T07:46:49.059192Z digest=sha256:916794f0b392a15acecf87a4a47c15b5806810e3c9bb4f0273166f85097f9be4

Observation e780d499-6b41-4c88-aca9-5acb343089f9 · inbound

Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos cites this paper.

Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos Listen, Think, and Understand

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-01T21:22:41.155465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T21:22:41.155465Z digest=sha256:6d102053b1a9af356c860420b1202fea02761a4ee505c87bfbe8ca8498176f47