Pith. sign in

Paper Citation Record · LEDGER

HASRD: Hierarchical Acoustic and Semantic Representation Disentanglement

As of 19 August 2026, this Paper Citation Record lists 37 of 37 outbound references and 3 inbound Pith citation observations for arXiv:2506.00843.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.00843 v1

Coverage vector

measured 37 of 37 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:01:43.372701Z

measured 40 of 40 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:01:43.249937Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T19:46:15.415472Z

Reference resolution

37 of 37 outbound references displayed

  • verified exact0
  • verified fuzzy24
  • unresolved13
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation db0b3957-29f6-4b6f-b8cd-57d4aaa09ccb · outbound

This paper cites an unresolved cited work.

HASRD: Hierarchical Acoustic and Semantic Representation Disentanglement Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:01:43.870474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T12:01:43.240637Z digest=sha256:aaf5905d937918df45c011bfdc667aa542630a6df943b4c44694a0b4ec1d33ab

Observation b4929345-79bb-4128-8beb-68d305901ca9 · outbound

This paper cites an unresolved cited work.

HASRD: Hierarchical Acoustic and Semantic Representation Disentanglement Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:01:43.860707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T12:01:43.244948Z digest=sha256:a23c9b7e4a4122d37697f5776bae5d660aa787b44c2604c6edcef296d1de76af

Observation 8bd3210c-2a8c-4609-9fee-b50842449518 · outbound

This paper cites HASRD: Hierarchical Acoustic and Semantic Representation Disentanglement.

HASRD: Hierarchical Acoustic and Semantic Representation Disentanglement HASRD: Hierarchical Acoustic and Semantic Representation Disentanglement

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T12:01:43.249937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:01:43.249937Z digest=sha256:da0dddff62ab3e3efa8a446ad8a80788110a3550b90659523a6528913c20d40d

Observation eeba0d67-d4b5-4d19-b4af-2ae36ba30a39 · outbound

This paper cites For acoustic training, we adopt the DAC framework [16], extracting random 5-second segments (vs.

HASRD: Hierarchical Acoustic and Semantic Representation Disentanglement For acoustic training, we adopt the DAC framework [16], extracting random 5-second segments (vs

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:43.851385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T12:01:43.253983Z digest=sha256:983da1a02b5cb07a1248b27a0686e57c396b97779a20d4266c52d5ea0455a678

Observation 82947716-31aa-43ff-bbdf-3dc0d1a48e89 · outbound

This paper cites an unresolved cited work.

HASRD: Hierarchical Acoustic and Semantic Representation Disentanglement Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:01:43.840244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T12:01:43.258313Z digest=sha256:bfabc11f796e76bbcb222b4eaa2fcabb182760fd7446a4b5c367ab0ba6b619dd

Observation 1e0b5dee-e1bc-4f0a-8c5a-9f05b98671d7 · outbound

This paper cites Our approach effectively preserves semantic performance for ASR while achieving reconstruction quality comparable to state-of-the-art neural audio codecs like DAC.

HASRD: Hierarchical Acoustic and Semantic Representation Disentanglement Our approach effectively preserves semantic performance for ASR while achieving reconstruction quality comparable to state-of-the-art neural audio codecs like DAC

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:43.828795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T12:01:43.262122Z digest=sha256:406fb789a621d41bc43bbfc31ece938e6cff4bf96ef09cb754cf8a86a9b875d7

Observation 0da11723-e0cd-4d2b-9188-999a55a0e457 · outbound

This paper cites Comparing Discrete and Continuous Space LLMs for Speech Recognition.

HASRD: Hierarchical Acoustic and Semantic Representation Disentanglement Comparing Discrete and Continuous Space LLMs for Speech Recognition

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T12:01:43.265309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:01:43.265309Z digest=sha256:6d3b69d0ed155c215839ffbe81e1c0ced02c2e5894fb2052583b628d3975fb2b

Observation aa2e66cf-500f-461a-a0b3-adc34cf2a18f · outbound

This paper cites AudioLM: A language modeling approach to audio generation,.

HASRD: Hierarchical Acoustic and Semantic Representation Disentanglement AudioLM: A language modeling approach to audio generation,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:43.816128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T12:01:43.268694Z digest=sha256:69b7ddf5fa95e3e2bad02b0a6e9f9a316c83cbfb219a94b478605050ab444aff

Observation 2693840c-85e6-414a-85aa-b7687e230ab6 · outbound

This paper cites A survey on speech large language models,.

HASRD: Hierarchical Acoustic and Semantic Representation Disentanglement A survey on speech large language models,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T12:01:43.272054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:01:43.272054Z digest=sha256:59d9e772c0f25edec9e55392bde38fadec319c40bf800faaef24e74522fbd36d

Observation 99b8b4a1-1f61-4892-aff3-7334e25a17f6 · outbound

This paper cites SpeechGPT: Empowering large language models with in- trinsic cross-modal conversational abilities,.

HASRD: Hierarchical Acoustic and Semantic Representation Disentanglement SpeechGPT: Empowering large language models with in- trinsic cross-modal conversational abilities,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:43.800601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T12:01:43.275346Z digest=sha256:8d9f30e836d2f510997a74fd4bb9d19f2f05d1980a903e1b8c5d729facf0515d

Observation 8267a6c4-ff08-4903-b453-4e3fed29d9db · outbound

This paper cites SpeechTokenizer: Unified Speech Tokenizer for Speech Large Language Models.

HASRD: Hierarchical Acoustic and Semantic Representation Disentanglement SpeechTokenizer: Unified Speech Tokenizer for Speech Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T12:01:43.278401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:01:43.278401Z digest=sha256:0c070ae49cc2347df5aa0f8ec4a8dda7fadb036546d0dce6b1c82c8ab4bfd042

Observation 7f951cb4-542c-4b3a-b069-28b70dc236dd · outbound

This paper cites Exploring speech recognition, translation, and understanding with discrete speech units: A comparative study,.

HASRD: Hierarchical Acoustic and Semantic Representation Disentanglement Exploring speech recognition, translation, and understanding with discrete speech units: A comparative study,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:43.788909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T12:01:43.281631Z digest=sha256:449f55659904f77b7329c0d5b63196e73933f2f2ec116af246e784e5cc1072e8

Observation a28f4ebe-6fac-4bd1-9202-ff7f9a27a104 · outbound

This paper cites HuBERT: Self-supervised speech representation learning by masked prediction of hidden units,.

HASRD: Hierarchical Acoustic and Semantic Representation Disentanglement HuBERT: Self-supervised speech representation learning by masked prediction of hidden units,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:43.779218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T12:01:43.285176Z digest=sha256:2e1e30c1043542ca39bfec43bc63ffde2fd5d8326e76009b2c87957dd73cc93c

Observation 62870af1-70b1-4e6b-add9-e5cae4802663 · outbound

This paper cites WavLM: Large-scale self- supervised pre-training for full stack speech processing,.

HASRD: Hierarchical Acoustic and Semantic Representation Disentanglement WavLM: Large-scale self- supervised pre-training for full stack speech processing,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:43.770201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T12:01:43.288493Z digest=sha256:c3c55c7ef2842a6a0a5b29451858cf52671263b0f7b0fd1d603f22999e6c97f8

Observation 39eac272-4f3c-464f-ab0c-9e0b666e396d · outbound

This paper cites W2v-BERT: Combining contrastive learning and masked language modeling for self-supervised speech pre- training,.

HASRD: Hierarchical Acoustic and Semantic Representation Disentanglement W2v-BERT: Combining contrastive learning and masked language modeling for self-supervised speech pre- training,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:43.761050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T12:01:43.293070Z digest=sha256:3ddc347766146969bdfb1063e123b1f65097a3bf503da463d1c852310eaa4736

Observation b7bc03b7-2eb0-481f-9ebb-67c954e0e0d9 · outbound

This paper cites Ex- ploration of efficient end-to-end ASR using discretized input from self-supervised learning,.

HASRD: Hierarchical Acoustic and Semantic Representation Disentanglement Ex- ploration of efficient end-to-end ASR using discretized input from self-supervised learning,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:43.750837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T12:01:43.297026Z digest=sha256:20fdbc6471c7d280d0fd75f8b43b5aabde1839ea3b9d0d564f044408db75473e

Observation 63762711-6ffb-4f12-9c90-beb58a970dd6 · outbound

This paper cites Speech resynthesis from discrete disentangled self-supervised representations,.

HASRD: Hierarchical Acoustic and Semantic Representation Disentanglement Speech resynthesis from discrete disentangled self-supervised representations,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:43.740820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T12:01:43.300050Z digest=sha256:55883356d3c4b03487227341e26053405afb498655d1ae09fdef8359fde6f091

Observation 1debd9fd-98c2-4a19-af30-efeb8f8a53c2 · outbound

This paper cites Ex- presso: A benchmark and analysis of discrete expressive speech resynthesis,.

HASRD: Hierarchical Acoustic and Semantic Representation Disentanglement Ex- presso: A benchmark and analysis of discrete expressive speech resynthesis,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:43.731052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T12:01:43.303897Z digest=sha256:33abc72de285c2954dbbdcfb2b3eb2f582fc75e03adb95927ee1e6540c6005e7

Observation ad690667-0ceb-4340-948d-a5881e8ed5dc · outbound

This paper cites SoundStream: An end-to-end neural audio codec,.

HASRD: Hierarchical Acoustic and Semantic Representation Disentanglement SoundStream: An end-to-end neural audio codec,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:43.720715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T12:01:43.307670Z digest=sha256:1cc5b635aac57507ac16317f84d486389c9b3eac58055e13e5403ac047593fb1

Observation 14e5435b-bbc6-4a45-84ca-5fc60c5ad807 · outbound

This paper cites High Fidelity Neural Audio Compression.

HASRD: Hierarchical Acoustic and Semantic Representation Disentanglement High Fidelity Neural Audio Compression

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T12:01:43.311743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:01:43.311743Z digest=sha256:03c303345c6a5284271104fd2341cb4858736afa09e2f4373b280565fbff2bf8

Observation 1bec4980-6a30-4924-a5a5-417b5b962cdc · outbound

This paper cites Moshi: a speech-text foundation model for real-time dialogue.

HASRD: Hierarchical Acoustic and Semantic Representation Disentanglement Moshi: a speech-text foundation model for real-time dialogue

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T12:01:43.315618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:01:43.315618Z digest=sha256:ba40f73fc64e70eaf0de1987042e5b077fdbaadb998c21e92323b6cccd16ed40

Observation 69d89b07-bf94-43c7-b1d9-5e91394365e4 · outbound

This paper cites High-fidelity audio compression with improved rvqgan,.

HASRD: Hierarchical Acoustic and Semantic Representation Disentanglement High-fidelity audio compression with improved rvqgan,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:43.708488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T12:01:43.319166Z digest=sha256:c1ac282f8d4d7a8f62a54a400b2644cc2e6a3eb1e543759a8e2c6025c2d74072

Observation 5b674017-f082-4303-b5d2-94a79dfa57b3 · outbound

This paper cites Codec Does Matter: Exploring the Semantic Shortcoming of Codec for Audio Language Model.

HASRD: Hierarchical Acoustic and Semantic Representation Disentanglement Codec Does Matter: Exploring the Semantic Shortcoming of Codec for Audio Language Model

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T12:01:43.322223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:01:43.322223Z digest=sha256:c4f413a3adde523feb806493580f5e956f20abf447fda47b6ebdaee4ea488a57

Observation 2c5d41c7-f072-4f97-8fc0-eb203deab3ff · outbound

This paper cites QR-VC: Leveraging Quantization Residuals for Linear Disentanglement in Zero-Shot Voice Conversion.

HASRD: Hierarchical Acoustic and Semantic Representation Disentanglement QR-VC: Leveraging Quantization Residuals for Linear Disentanglement in Zero-Shot Voice Conversion

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T12:01:43.325585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:01:43.325585Z digest=sha256:0fcda5a67711b544c2523a8f3f2a8457fe809e973534b36f9611fb955276aafd

Observation ac8b6874-1279-488e-ab19-fcd25f5d331f · outbound

This paper cites MMM: Multi-layer multi-residual multi-stream discrete speech repre- sentation from self-supervised learning model,.

HASRD: Hierarchical Acoustic and Semantic Representation Disentanglement MMM: Multi-layer multi-residual multi-stream discrete speech repre- sentation from self-supervised learning model,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:43.697823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T12:01:43.329456Z digest=sha256:8081a966426ffb72d479e78640610eef6a8944111e770fba00869309b4233eda

Observation 9550b58b-0ce6-43f0-9933-47545d4125e1 · outbound

This paper cites Towards universal speech discrete tokens: A case study for ASR and TTS,.

HASRD: Hierarchical Acoustic and Semantic Representation Disentanglement Towards universal speech discrete tokens: A case study for ASR and TTS,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:43.687447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T12:01:43.332570Z digest=sha256:c334b23259ffc34f8497799a1203350d88fb90b56625b2585acc56d47b26866c

Observation 5229250d-6dd1-4705-8258-566427df1879 · outbound

This paper cites ReVISE: Self-supervised speech resynthesis with visual input for univer- sal and generalized speech regeneration,.

HASRD: Hierarchical Acoustic and Semantic Representation Disentanglement ReVISE: Self-supervised speech resynthesis with visual input for univer- sal and generalized speech regeneration,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:43.675974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T12:01:43.335931Z digest=sha256:956a4a77985af434deeb8f9586180c5101d70af45dedd592ff5f4441c898c154

Observation c93f827d-aab7-45c8-ab08-d67d4ec6893a · outbound

This paper cites Self-supervised disentan- gled representation learning for robust target speech extraction,.

HASRD: Hierarchical Acoustic and Semantic Representation Disentanglement Self-supervised disentan- gled representation learning for robust target speech extraction,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:43.664679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T12:01:43.339284Z digest=sha256:2dd0306fd330483b2b46007cc6c17f992c29b537f46f72bb833ac51314adcaa4

Observation 5457c13d-fa21-4af3-aece-434aff075c76 · outbound

This paper cites A ConvNet for the 2020s,.

HASRD: Hierarchical Acoustic and Semantic Representation Disentanglement A ConvNet for the 2020s,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:43.653509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T12:01:43.342683Z digest=sha256:cc13348111f1732b328e1f4e3ec6f9f5e3cb84f8bf5ee81dc9211b1112dd11a2

Observation 1c0dc1e9-17c3-4a52-841d-b2ddf644724c · outbound

This paper cites Self-supervised learning with random-projection quantizer for speech recogni- tion,.

HASRD: Hierarchical Acoustic and Semantic Representation Disentanglement Self-supervised learning with random-projection quantizer for speech recogni- tion,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:43.643567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T12:01:43.345879Z digest=sha256:b83d3e023c1279b67c8e676d5d3de9c8f3166eade722fa0f17e7a91de6c58f09

Observation 7cbdc036-e0b1-4139-b856-a6d2a7927a1a · outbound

This paper cites Mel- GAN: Generative adversarial networks for conditional wave- form synthesis,.

HASRD: Hierarchical Acoustic and Semantic Representation Disentanglement Mel- GAN: Generative adversarial networks for conditional wave- form synthesis,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:43.633237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T12:01:43.349431Z digest=sha256:777181bb283b9bec643c7c9347b18347c329b2a08a5861b355b09d53166bbf85

Observation 46499834-a384-4020-954c-719ab49dde7d · outbound

This paper cites HiFi-GAN: Generative adversarial networks for efficient and high fidelity speech synthesis,.

HASRD: Hierarchical Acoustic and Semantic Representation Disentanglement HiFi-GAN: Generative adversarial networks for efficient and high fidelity speech synthesis,

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T12:01:43.353811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:01:43.353811Z digest=sha256:16f5cfceb0dddc63a6df8a697da694cc72dedae638d824b05ce10a1624ccfca5

Observation 8f28f306-a4f7-4dee-a2c4-9d2e8bdd0af8 · outbound

This paper cites Lib- riSpeech: An ASR corpus based on public domain audio books,.

HASRD: Hierarchical Acoustic and Semantic Representation Disentanglement Lib- riSpeech: An ASR corpus based on public domain audio books,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:43.615831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T12:01:43.357062Z digest=sha256:a09dc832a7bee1c89b85339946b33b7348f9dfec6d8a66050288b8b5292005e5

Observation 13b65055-9954-4a90-baa7-7cd5312b1987 · outbound

This paper cites Lhotse: A speech data representation library for the modern deep learning ecosystem,.

HASRD: Hierarchical Acoustic and Semantic Representation Disentanglement Lhotse: A speech data representation library for the modern deep learning ecosystem,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:43.605540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T12:01:43.360747Z digest=sha256:dfe5189077fd4f5ff54b358b355a6c44314e5bf884353e3de866c6a869cdda61

Observation e29dcf14-9587-42e5-ac29-b1656da52e00 · outbound

This paper cites Open Implementation and Study of BEST-RQ for Speech Processing.

HASRD: Hierarchical Acoustic and Semantic Representation Disentanglement Open Implementation and Study of BEST-RQ for Speech Processing

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T12:01:43.364043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:01:43.364043Z digest=sha256:3cc5e445048a59cd8862437bd3593c538793b22c5aa35cbe740fe8c821c7fe91

Observation a8a9fc1d-9240-4912-9dd7-2837dee109f3 · outbound

This paper cites Con- nectionist temporal classification: Labelling unsegmented se- quence data with recurrent neural networks,.

HASRD: Hierarchical Acoustic and Semantic Representation Disentanglement Con- nectionist temporal classification: Labelling unsegmented se- quence data with recurrent neural networks,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:43.595072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T12:01:43.368593Z digest=sha256:91ab196dd28e8c1386da1bf7a68f76d63277b91d97bf69a79102698ad20cf4c7

Observation d64d0e91-ea06-4e9c-87b0-c57f1cf42165 · outbound

This paper cites ECAPA- TDNN: Emphasized channel attention, propagation and aggre- gation in TDNN based speaker verification,.

HASRD: Hierarchical Acoustic and Semantic Representation Disentanglement ECAPA- TDNN: Emphasized channel attention, propagation and aggre- gation in TDNN based speaker verification,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:43.584006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T12:01:43.372701Z digest=sha256:63d3fc90a78abdeac9b423996000db682bc69d48612a42b2775d50faf177af1f

Pith citing papers

Observation 8bd3210c-2a8c-4609-9fee-b50842449518 · inbound

HASRD: Hierarchical Acoustic and Semantic Representation Disentanglement cites this paper.

HASRD: Hierarchical Acoustic and Semantic Representation Disentanglement HASRD: Hierarchical Acoustic and Semantic Representation Disentanglement

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T12:01:43.249937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:01:43.249937Z digest=sha256:da0dddff62ab3e3efa8a446ad8a80788110a3550b90659523a6528913c20d40d

Observation da8f0ff8-c3eb-4a88-8fbf-6acdd6e5f6ab · inbound

Minimizing Modality Gap from the Input Side: Your Speech LLM Can Be a Prosody-Aware Text LLM cites this paper.

Minimizing Modality Gap from the Input Side: Your Speech LLM Can Be a Prosody-Aware Text LLM HASRD: Hierarchical Acoustic and Semantic Representation Disentanglement

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:46:15.417161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-08T11:00:52.196039Z digest=sha256:d07398f6c164c7e72bc393c18ebfb5d2f4b8b26d99cf14c6b673838ae79e1ed1

Observation 98a9678c-94cd-430a-88af-17b774c73b52 · inbound

Minimizing Modality Gap from the Input Side: Your Speech LLM Can Be a Prosody-Aware Text LLM cites this paper.

Minimizing Modality Gap from the Input Side: Your Speech LLM Can Be a Prosody-Aware Text LLM HASRD: Hierarchical Acoustic and Semantic Representation Disentanglement

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:50:49.653335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-11T00:49:26.507281Z digest=sha256:447546c64b8d7e9d42b747fb57592e85c0ba545b03ed97d788af4c63339165d3