Pith. sign in

Paper Citation Record · LEDGER

Listen, Do Not Copy: Internalizing Audio-Grounded Scaffold Context for Robust Omni-Model Speech Understanding

As of 9 August 2026, this Paper Citation Record lists 20 of 20 outbound references and 0 inbound Pith citation observations for arXiv:2607.21943.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.21943 v2

Coverage vector

measured 20 of 20 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-01T06:19:33.926800Z

measured 20 of 20 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

20 of 20 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved20
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation bac90fb4-1169-4c14-806f-bcf450708ada · outbound

This paper cites Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models.

Listen, Do Not Copy: Internalizing Audio-Grounded Scaffold Context for Robust Omni-Model Speech Understanding Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-01T06:19:31.491739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:19:31.491739Z digest=sha256:c6ace568149f8e4968a2f799c44793c2a611c6eb5feef0813c172457b5d9a18c

Observation 2329e598-3981-4897-a512-67570bec96af · outbound

This paper cites DOI:10.1038/s42256-020-00257-z.

Listen, Do Not Copy: Internalizing Audio-Grounded Scaffold Context for Robust Omni-Model Speech Understanding DOI:10.1038/s42256-020-00257-z

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-01T06:19:32.061870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:19:32.061870Z digest=sha256:6bf37a2a00d5f6458f1f7eb7379df5d73fcae76097d4a859f972033c1ace7fa0

Observation f9308c80-372d-436c-85da-83b247d367a0 · outbound

This paper cites On the Audio Hallucinations in Large Audio-Video Language Models.

Listen, Do Not Copy: Internalizing Audio-Grounded Scaffold Context for Robust Omni-Model Speech Understanding On the Audio Hallucinations in Large Audio-Video Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-01T06:19:32.544745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:19:32.544745Z digest=sha256:c268cb5e78b57a24118cfb2a718f4a14a06c1dc6cd08830d72f86e71e782e998

Observation 81c7a2e4-75f7-4b91-926b-24a9a246e60b · outbound

This paper cites Proximal Policy Optimization Algorithms.

Listen, Do Not Copy: Internalizing Audio-Grounded Scaffold Context for Robust Omni-Model Speech Understanding Proximal Policy Optimization Algorithms

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T06:19:32.664754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:19:32.664754Z digest=sha256:7966139ac9383355eb3aef567d16a98c4d24763ed068f88219290b137140ef22

Observation e6c8895e-7140-4b0d-9c91-a653f644cb73 · outbound

This paper cites Learning by Distilling Context.

Listen, Do Not Copy: Internalizing Audio-Grounded Scaffold Context for Robust Omni-Model Speech Understanding Learning by Distilling Context

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-01T06:19:33.044773Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:19:33.044773Z digest=sha256:9530c1603032d3e49076ca4005c3eba9a5a55938ad8609ff88799e36568a1f56

Observation 1d7d3649-75ec-4377-8aa6-2bdb4a079cc4 · outbound

This paper cites InProceedings of the Annual Conference of the International Speech Commu- nication Association (INTERSPEECH), 3754–3758.

Listen, Do Not Copy: Internalizing Audio-Grounded Scaffold Context for Robust Omni-Model Speech Understanding InProceedings of the Annual Conference of the International Speech Commu- nication Association (INTERSPEECH), 3754–3758

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-01T06:19:33.564755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:19:33.564755Z digest=sha256:75de1c507340c993fb360bdd1d320122aa735896be9d3777e9cb7dc7a77db9c9

Observation 7b596cc1-0481-4615-bc25-834de344f142 · outbound

This paper cites Qwen3-Omni Technical Report.

Listen, Do Not Copy: Internalizing Audio-Grounded Scaffold Context for Robust Omni-Model Speech Understanding Qwen3-Omni Technical Report

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T06:19:33.661907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:19:33.661907Z digest=sha256:56e7437dff37a25b2ea0ee49e34a8fc9ab4c347f8dcbbf1d3fa30ae509486d66

Observation 4e4da53f-4006-499d-842e-e4785c06da0f · outbound

This paper cites InProceedings of the Annual Meeting of the Association for Computational Linguistics (ACL), 1979–1998.

Listen, Do Not Copy: Internalizing Audio-Grounded Scaffold Context for Robust Omni-Model Speech Understanding InProceedings of the Annual Meeting of the Association for Computational Linguistics (ACL), 1979–1998

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-01T06:19:33.821036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:19:33.821036Z digest=sha256:395ca051a69160c04412c1695ea7c9d1d9c1d8173fa7c7c3997c8e56fc1e7542

Observation 6bc1d1a8-3aa3-4c73-9b94-7aa64cc1da89 · outbound

This paper cites IRAF: Interference-Resilient Adaptive Fusion for Noise-Robust End-to-End Full-Duplex Spoken Dialogue Systems.

Listen, Do Not Copy: Internalizing Audio-Grounded Scaffold Context for Robust Omni-Model Speech Understanding IRAF: Interference-Resilient Adaptive Fusion for Noise-Robust End-to-End Full-Duplex Spoken Dialogue Systems

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-01T06:19:33.926800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:19:33.926800Z digest=sha256:b2d5b931e77485217fddba4b05ba0983f3e84469b23d0a2d57429ccb8f41c4a7

Observation bfb6310e-03db-464e-92ca-9c3b95e43170 · outbound

This paper cites InProceedings of the International Conference on Learning Representations (ICLR), volume 2024, 16607– 16629.

Listen, Do Not Copy: Internalizing Audio-Grounded Scaffold Context for Robust Omni-Model Speech Understanding InProceedings of the International Conference on Learning Representations (ICLR), volume 2024, 16607– 16629

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-01T06:19:33.354005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:19:33.354005Z digest=sha256:2bff9545f5947f444e59b8525aa4de2887b5678dc6b10c0f62027210da981645

Observation 6e19b0ee-178e-4a38-b9c1-e4147a28eb0f · outbound

This paper cites Springer.

Listen, Do Not Copy: Internalizing Audio-Grounded Scaffold Context for Robust Omni-Model Speech Understanding Springer

Reference 2006

Resolution
unresolved
no resolver link, observed 2026-08-01T06:19:31.119543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:19:31.119543Z digest=sha256:db54c5fac4d35a6b7831dc9c20bf1ce76c36297c2e4878bab24ab31896d6c917

Observation 4023d0a1-24ef-4957-8522-b17596bca7b9 · outbound

This paper cites MUSAN: A Music, Speech, and Noise Corpus.

Listen, Do Not Copy: Internalizing Audio-Grounded Scaffold Context for Robust Omni-Model Speech Understanding MUSAN: A Music, Speech, and Noise Corpus

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-01T06:19:33.244754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:19:33.244754Z digest=sha256:1903a2c339bb243a108ce51c4e6b1b75218e1db7371aff774bc7df691bde841f

Observation 441e5737-9cdc-4e86-836a-f6ced4cd9c14 · outbound

This paper cites In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 6904–6913.

Listen, Do Not Copy: Internalizing Audio-Grounded Scaffold Context for Robust Omni-Model Speech Understanding In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 6904–6913

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-01T06:19:32.186964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:19:32.186964Z digest=sha256:f7ae4c96a897d4ca7eaebba868f537f9e472340e74b03ca8e4a19f07de4041d3

Observation 432e0dbd-1837-496a-b54b-a55e0dda9d31 · outbound

This paper cites LibriMix: An Open-Source Dataset for Generalizable Speech Separation.

Listen, Do Not Copy: Internalizing Audio-Grounded Scaffold Context for Robust Omni-Model Speech Understanding LibriMix: An Open-Source Dataset for Generalizable Speech Separation

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-01T06:19:31.674770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:19:31.674770Z digest=sha256:7990000180a7f083c791c2eea4d27919766fa6dcac04911fd9e0f312cc63e232

Observation 7b6d27cd-4300-4469-af01-819a2847d16e · outbound

This paper cites InProceedings of the Annual Conference of the In- ternational Speech Communication Association (INTER- SPEECH), 2756–2760.

Listen, Do Not Copy: Internalizing Audio-Grounded Scaffold Context for Robust Omni-Model Speech Understanding InProceedings of the Annual Conference of the In- ternational Speech Communication Association (INTER- SPEECH), 2756–2760

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-01T06:19:32.911504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:19:32.911504Z digest=sha256:cd302097a3513d5026ef5d3b03701e54f43d9a08dcc8a171d4898dbcce0abe58

Observation 8dec6793-8112-4bb7-9124-d22c4c7b609b · outbound

This paper cites DOI: 10.1109/JSTSP.2022.3188113.

Listen, Do Not Copy: Internalizing Audio-Grounded Scaffold Context for Robust Omni-Model Speech Understanding DOI: 10.1109/JSTSP.2022.3188113

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-01T06:19:31.334755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:19:31.334755Z digest=sha256:387603508960167b0ed5ca18fced634ef080f261941e7aa9947b338a34eda728

Observation 32b2485b-8812-424b-a8d5-53c300c1bb1d · outbound

This paper cites InProceedings of the Annual Conference of the International Speech Com- munication Association (INTERSPEECH), 1983–1987.

Listen, Do Not Copy: Internalizing Audio-Grounded Scaffold Context for Robust Omni-Model Speech Understanding InProceedings of the Annual Conference of the International Speech Com- munication Association (INTERSPEECH), 1983–1987

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-01T06:19:30.987089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:19:30.987089Z digest=sha256:2f31ee3847dc04f5da4959475a2ad870346ff7f40dfeb2ed97067bc3e9efee29

Observation b7d6eea6-229f-466a-ab8f-780b4ceeee89 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Listen, Do Not Copy: Internalizing Audio-Grounded Scaffold Context for Robust Omni-Model Speech Understanding DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-01T06:19:32.754751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:19:32.754751Z digest=sha256:eb3644546f75318eb2a353f312c9111b7c1fd7d6c2d9ab220aa3595d78fa8010

Observation 3ec5dbe5-adf1-4afc-9851-636b61d723ae · outbound

This paper cites arXiv:2510.24821.

Listen, Do Not Copy: Internalizing Audio-Grounded Scaffold Context for Robust Omni-Model Speech Understanding arXiv:2510.24821

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-01T06:19:32.354752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:19:32.354752Z digest=sha256:a631dc2f54ccf689730db136db1e5c5328e0dba4527fd26a51c54e4bdd111763

Observation 836bde2c-c03e-4656-86fd-ff08ff939653 · outbound

This paper cites MiniCPM-o 4.5: Towards Real-Time Full-Duplex Omni-Modal Interaction.

Listen, Do Not Copy: Internalizing Audio-Grounded Scaffold Context for Robust Omni-Model Speech Understanding MiniCPM-o 4.5: Towards Real-Time Full-Duplex Omni-Modal Interaction

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-01T06:19:31.834213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:19:31.834213Z digest=sha256:207a9881798eda8b2cc7ba50dc1b4541ad28515e5d4c88d2f239d385c94044ec

Pith citing papers

No inbound Pith citation observations are available.