Pith. sign in

Paper Citation Record · LEDGER

Hear, Invoke, and Understand: A Skill-Calling Multimodal Agent for Large Audio Language Models

As of 20 August 2026, this Paper Citation Record lists 25 of 25 outbound references and 0 inbound Pith citation observations for arXiv:2608.01881.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.01881 v1

Coverage vector

measured 25 of 25 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T19:02:41.528122Z

measured 25 of 25 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

25 of 25 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved25
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 77fb404d-c569-4f15-a722-967f86973136 · outbound

This paper cites FSD50K: An Open Dataset of Human-Labeled Sound Events.

Hear, Invoke, and Understand: A Skill-Calling Multimodal Agent for Large Audio Language Models FSD50K: An Open Dataset of Human-Labeled Sound Events

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T19:02:40.771164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:02:40.771164Z digest=sha256:4e304928366b4fedbc743a63f925dc1a6400dfee97d43faf536f54596718e82d

Observation 7211c381-a8d4-40ec-9e22-b8567e9480d9 · outbound

This paper cites Full-Duplex-Bench-v3: Benchmarking Tool Use for Full-Duplex Voice Agents Under Real-World Disfluency.

Hear, Invoke, and Understand: A Skill-Calling Multimodal Agent for Large Audio Language Models Full-Duplex-Bench-v3: Benchmarking Tool Use for Full-Duplex Voice Agents Under Real-World Disfluency

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T19:02:40.932166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:02:40.932166Z digest=sha256:2311393ac806b06b224894220bb0214c452ef74ab92158494d2046973a2c2b6b

Observation 2e2a3dbc-9498-453c-a23b-34a3e2ccf11a · outbound

This paper cites MMAR: A Challenging Benchmark for Deep Reasoning in Speech, Audio, Music, and Their Mix.

Hear, Invoke, and Understand: A Skill-Calling Multimodal Agent for Large Audio Language Models MMAR: A Challenging Benchmark for Deep Reasoning in Speech, Audio, Music, and Their Mix

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T19:02:41.013510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:02:41.013510Z digest=sha256:746be03a579d3c26efd27411ab9ab024505087a3e43350e473fff0dc17d4672c

Observation 6187c0fc-5fe1-4a16-ad07-e5a4d9c68382 · outbound

This paper cites InIEEE Au- tomatic Speech Recognition and Understanding Workshop, ASRU2025,Honolulu,HI,USA,December6-10,2025,1–4.

Hear, Invoke, and Understand: A Skill-Calling Multimodal Agent for Large Audio Language Models InIEEE Au- tomatic Speech Recognition and Understanding Workshop, ASRU2025,Honolulu,HI,USA,December6-10,2025,1–4

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T19:02:41.053261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:02:41.053261Z digest=sha256:0b90a876e89d2c947f2dd587737a00f64585377ea6c5480aa8c67f6169e4e0d7

Observation 8b09f8e2-b07b-4cbf-ad50-41d679f06ce7 · outbound

This paper cites In Lacerda, F., ed.,18th Annual Conference of the Inter- national Speech Communication Association, Interspeech 2017, Stockholm, Sweden, August 20-24, 2017, 2616–2620.

Hear, Invoke, and Understand: A Skill-Calling Multimodal Agent for Large Audio Language Models In Lacerda, F., ed.,18th Annual Conference of the Inter- national Speech Communication Association, Interspeech 2017, Stockholm, Sweden, August 20-24, 2017, 2616–2620

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T19:02:41.067683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:02:41.067683Z digest=sha256:7449e920042e35b2c20a77115f9a6b75d7b8d7ca95bd364729c4c234f2424b99

Observation 34f5ce39-792b-48b8-87a2-61e208bc2c30 · outbound

This paper cites Audio2Tool: Speak, Call, Act -- A Dataset for Benchmarking Speech Tool Use.

Hear, Invoke, and Understand: A Skill-Calling Multimodal Agent for Large Audio Language Models Audio2Tool: Speak, Call, Act -- A Dataset for Benchmarking Speech Tool Use

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T19:02:41.107158Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:02:41.107158Z digest=sha256:bf95d676997c34d7cd01cbc4f714d734a22bf6daf32fb852c1604577c08b34e5

Observation 1b4f6fc5-f7b4-4187-9e5a-c35004e175f5 · outbound

This paper cites Sakshi, S.; Tyagi, U.; and Kumar, S.

Hear, Invoke, and Understand: A Skill-Calling Multimodal Agent for Large Audio Language Models Sakshi, S.; Tyagi, U.; and Kumar, S

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T19:02:41.148943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:02:41.148943Z digest=sha256:54cfca6d3b398a0b8de19a492dac52b931a1a441d5197058e0cb35914a70030e

Observation 98553ed0-0e79-4d3e-a5ce-dc733f5d1faa · outbound

This paper cites Qwen3-Omni Technical Report.

Hear, Invoke, and Understand: A Skill-Calling Multimodal Agent for Large Audio Language Models Qwen3-Omni Technical Report

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T19:02:41.221713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:02:41.221713Z digest=sha256:4f4e4ba470a3da1585fd0fdb25b9442bbdbf6cfe2b075c653dd6c0993ac7dc21

Observation bc62cbbc-a02d-4d4d-bb71-bace030c9f48 · outbound

This paper cites Qwen3.5-Omni Technical Report.

Hear, Invoke, and Understand: A Skill-Calling Multimodal Agent for Large Audio Language Models Qwen3.5-Omni Technical Report

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T19:02:41.256146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:02:41.256146Z digest=sha256:684e3d20f4d14115d7bcf65256439809f35784ddc180ff8dc2549279be75a0cc

Observation 4b09b3f9-95d8-4749-8218-152e34f768e0 · outbound

This paper cites Tong, S.; Li, X.; and Wang, Y.

Hear, Invoke, and Understand: A Skill-Calling Multimodal Agent for Large Audio Language Models Tong, S.; Li, X.; and Wang, Y

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T19:02:41.293091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:02:41.293091Z digest=sha256:1b5df1d7ca87b620fd4c2c3e83435b9d611ce2c58c724e5d75e0357cc0afce6a

Observation 3f03065c-a706-42ed-8410-2f5b6d2d8c21 · outbound

This paper cites Wang, B.; Zou, X.; and Lin, G.

Hear, Invoke, and Understand: A Skill-Calling Multimodal Agent for Large Audio Language Models Wang, B.; Zou, X.; and Lin, G

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T19:02:41.335720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:02:41.335720Z digest=sha256:87d42ad2cd393d92e8de727e2332aa3bd7417e2a1e1e6f117e361f84fc2c6e19

Observation 6f0e38b8-5a61-44bb-a818-8123a40a4ac6 · outbound

This paper cites an unresolved cited work.

Hear, Invoke, and Understand: A Skill-Calling Multimodal Agent for Large Audio Language Models Unresolved cited work

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T19:02:41.353060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:02:41.353060Z digest=sha256:d2166a10f37f59ac3ffaac9c94006b2e0e41a2e1365e7cdd9c7794846c5f3a53

Observation f91bf7a1-04a2-4871-b173-21ed1dd452e9 · outbound

This paper cites MSU-Bench: Towards Understanding the Conversational Multi-talker Scenarios.

Hear, Invoke, and Understand: A Skill-Calling Multimodal Agent for Large Audio Language Models MSU-Bench: Towards Understanding the Conversational Multi-talker Scenarios

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-04T19:02:41.392890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:02:41.392890Z digest=sha256:99b456893c19686abe367a814bad4ae64d8c332b55f3d4dfacecbf18e8d8370d

Observation c834345b-e7e9-450f-a683-d09f2176cb2e · outbound

This paper cites Audio-Mind: An Auditable Agentic Framework for Audio Understanding.

Hear, Invoke, and Understand: A Skill-Calling Multimodal Agent for Large Audio Language Models Audio-Mind: An Auditable Agentic Framework for Audio Understanding

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T19:02:41.428446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:02:41.428446Z digest=sha256:ede23f10247b81898370ab03404bd488824505d0ad521eb64fad77714d03d59d

Observation eb13be1e-37ad-453a-a48e-e22c2db5df73 · outbound

This paper cites Xie,Z.;Lin,M.;andLiu,Z.2025.

Hear, Invoke, and Understand: A Skill-Calling Multimodal Agent for Large Audio Language Models Xie,Z.;Lin,M.;andLiu,Z.2025

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-04T19:02:41.470759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:02:41.470759Z digest=sha256:c7ec0561c4861a5b51eee8fd0a97661ba2acb6b6c6d80345c5c4f1b808c72a5f

Observation 9be4a0b8-7f58-41f1-a43b-7f2bb47ef9ac · outbound

This paper cites CoRR, abs/2606.15141.

Hear, Invoke, and Understand: A Skill-Calling Multimodal Agent for Large Audio Language Models CoRR, abs/2606.15141

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T19:02:41.491746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:02:41.491746Z digest=sha256:a597bb85bdfd7a7f2e3d02d0b6a90caf0977f6e41b1cbc0208890bb4bb6ec4a2

Observation 13256ed0-1e9f-4e28-bbfc-217dd82fbb6a · outbound

This paper cites AISHELL-1: An Open-Source Mandarin Speech Corpus and A Speech Recognition Baseline.

Hear, Invoke, and Understand: A Skill-Calling Multimodal Agent for Large Audio Language Models AISHELL-1: An Open-Source Mandarin Speech Corpus and A Speech Recognition Baseline

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-04T19:02:40.653124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:02:40.653124Z digest=sha256:f100ee8eb4aef27a58850e5df4e5458bdbbfb856d8dcc560149bb5b7dd6334e5

Observation c8bcbcec-dc56-4ace-9a1a-ea4a8567596b · outbound

This paper cites Spoken SQuAD: A Study of Mitigating the Impact of Speech Recognition Errors on Listening Comprehension.

Hear, Invoke, and Understand: A Skill-Calling Multimodal Agent for Large Audio Language Models Spoken SQuAD: A Study of Mitigating the Impact of Speech Recognition Errors on Listening Comprehension

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-04T19:02:40.885771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:02:40.885771Z digest=sha256:e919e75d53754fa246448a2fe81dfb600e40e679ed4a9a5d79a8ef9b5bbc6270

Observation d71e7252-bb08-46c4-9f06-f0a42d6c960f · outbound

This paper cites LibriMix: An Open-Source Dataset for Generalizable Speech Separation.

Hear, Invoke, and Understand: A Skill-Calling Multimodal Agent for Large Audio Language Models LibriMix: An Open-Source Dataset for Generalizable Speech Separation

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-04T19:02:40.730450Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:02:40.730450Z digest=sha256:c1552d8b11df08a4276bd05b8ed6359477b694bc3bcb253cbfc55b6391be82d6

Observation 724f824a-7e93-4787-920e-4654dc3fc791 · outbound

This paper cites AISHELL-4: An Open Source Dataset for Speech Enhancement, Separation, Recognition and Speaker Diarization in Conference Scenario.

Hear, Invoke, and Understand: A Skill-Calling Multimodal Agent for Large Audio Language Models AISHELL-4: An Open Source Dataset for Speech Enhancement, Separation, Recognition and Speaker Diarization in Conference Scenario

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-04T19:02:40.804727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:02:40.804727Z digest=sha256:577c5d8052e02642dc331f087af68fed393fe463b0138b7d21d224958ec936fa

Observation 4dcc0451-20e1-44a7-b907-4a2f4d95a0fc · outbound

This paper cites ASVspoof 2021: Towards Spoofed and Deepfake Speech Detection in the Wild.

Hear, Invoke, and Understand: A Skill-Calling Multimodal Agent for Large Audio Language Models ASVspoof 2021: Towards Spoofed and Deepfake Speech Detection in the Wild

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-04T19:02:40.969960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:02:40.969960Z digest=sha256:40720adec47c885d5a76bff77968fb4ce0152c54667b744e66ea89dc923445b6

Observation 9c752120-f0a1-4d5f-b55f-f14617d2bbb0 · outbound

This paper cites LibriSQA: A Novel Dataset and Framework for Spoken Question Answering with Large Language Models.

Hear, Invoke, and Understand: A Skill-Calling Multimodal Agent for Large Audio Language Models LibriSQA: A Novel Dataset and Framework for Spoken Question Answering with Large Language Models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-04T19:02:41.528122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:02:41.528122Z digest=sha256:0af091c96b0ebf6dbb1eee5118b0f0ef2be085c062b65b397f050da6b3985709

Observation 132df9a7-37fb-47bd-a385-bfcccba9b1e1 · outbound

This paper cites MMAU: A Massive Multi-Task Audio Understanding and Reasoning Benchmark.

Hear, Invoke, and Understand: A Skill-Calling Multimodal Agent for Large Audio Language Models MMAU: A Massive Multi-Task Audio Understanding and Reasoning Benchmark

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-04T19:02:41.181470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:02:41.181470Z digest=sha256:412ba379df848d76a4d595631220b400e293838ee98f48fba1f6e8fd9b426b1d

Observation adbdab8b-0935-4d1e-abf1-c57a5fb4f5e9 · outbound

This paper cites KimiTeam;Ding,D.;andJu,Z.2025.

Hear, Invoke, and Understand: A Skill-Calling Multimodal Agent for Large Audio Language Models KimiTeam;Ding,D.;andJu,Z.2025

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-04T19:02:40.849020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:02:40.849020Z digest=sha256:63f72fb4b836deee56668d454f5be945830fbecc266405a9d0bbabb8974e95e3

Observation a4d5a1be-9bd3-4ea1-8cc1-7032541fc61f · outbound

This paper cites CoRR, abs/2602.10439.

Hear, Invoke, and Understand: A Skill-Calling Multimodal Agent for Large Audio Language Models CoRR, abs/2602.10439

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-04T19:02:40.693596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:02:40.693596Z digest=sha256:8d0499e765adcabae7f30b2b9a019a04fa7b0e96962304da345abaf0ac81e1b8

Pith citing papers

No inbound Pith citation observations are available.