Pith. sign in

Paper Citation Record · LEDGER

Breaking the Barriers of Text-Hungry and Audio-Deficient AI

As of 9 August 2026, this Paper Citation Record lists 100 of 165 outbound references and 3 inbound Pith citation observations for arXiv:2506.02443.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.02443 v1

Coverage vector

measured 100 of 165 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:28:54.976320Z

measured 103 of 103 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-26T09:46:15.701501Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T09:39:46.872455Z

Reference resolution

100 of 165 outbound references displayed

  • verified exact17
  • verified fuzzy0
  • unresolved83
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c6c2304a-23ac-4914-ba28-4f55fc9e71c3 · outbound

This paper cites Simultron: On-device simultaneous speech to speech translation.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Simultron: On-device simultaneous speech to speech translation

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:45.767223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:45.767223Z digest=sha256:6ec5602cfbb1f386884431cee528f2682e273c30177a19031ddfc51bdfe379fd

Observation 873835c3-1a3e-4f91-b08d-eda15d57db18 · outbound

This paper cites Qwen 2.5: A comprehensive review of the leading resource-efficient llm with potentioal to surpass all competitors.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Qwen 2.5: A comprehensive review of the leading resource-efficient llm with potentioal to surpass all competitors

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:45.821059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:45.821059Z digest=sha256:0e5c53f317e8e95bd21a1cf8b72f8f09be23e96e2822946ad7d985c9ee4e695a

Observation 32a0ea4b-8888-4f91-83a8-5d35e1858e85 · outbound

This paper cites Analysis of layer- wise training in direct speech to speech translation using bi-lstm.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Analysis of layer- wise training in direct speech to speech translation using bi-lstm

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:45.924856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:45.924856Z digest=sha256:5d38a633d4446e9f6dc28855d727125c85105bf4ce14f20918380e3a30176c7c

Observation 2a273406-4bde-489e-bdd9-71d54fcd955a · outbound

This paper cites Precipitation nowcasting with generative diffusion models.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Precipitation nowcasting with generative diffusion models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:46.038879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:46.038879Z digest=sha256:4c0c7fd096e1b2d2061b8b6b06fae55d1ae08fb3fe7e82577c7ef0caec2cb085

Observation 3b8fe074-38c4-4a16-aaab-be8e14324918 · outbound

This paper cites Mean-Field-Type Game Theory: Applica- tions, volume 2.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Mean-Field-Type Game Theory: Applica- tions, volume 2

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:46.191304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:46.191304Z digest=sha256:639924bb2700b14a8212a2c1dd32d36b0f26563d0c3fe1ae2a29160911a73425

Observation 50a12541-1011-418a-890b-61774aaf52b1 · outbound

This paper cites an unresolved cited work.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Unresolved cited work

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:46.297443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:46.297443Z digest=sha256:a37301d171eb868895babc0d7b3b411a8c5484ffea878a3490cd153fc4efc132

Observation f7990542-4be9-4a87-90b4-55da5b2b7ea4 · outbound

This paper cites Listen and Translate: A Proof of Concept for End-to-End Speech-to-Text Translation.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Listen and Translate: A Proof of Concept for End-to-End Speech-to-Text Translation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:46.397405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:46.397405Z digest=sha256:286c13e3e10e80a4e5bffc544c7d89df52241675e51c22deaa57cabead068417

Observation 1c201fa9-12a9-4fbd-933a-bcdce127e9ac · outbound

This paper cites an unresolved cited work.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Unresolved cited work

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:46.487794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:46.487794Z digest=sha256:88428f670b4862679c924917c98cb144963c690e260a60770002b1a377c8adde

Observation c72b7324-91de-4b01-9e1e-f8fae1dd1ebd · outbound

This paper cites Large language models are strong audio-visual speech recognition learners.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Large language models are strong audio-visual speech recognition learners

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:46.586257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:46.586257Z digest=sha256:78ceac7c2fd3e114c1e04515e418b7ae265f48892678f169d3d817bd7073d963

Observation 82501d3c-2e97-4271-9597-0295d15f6637 · outbound

This paper cites Low frame-rate speech codec: a codec designed for fast high-quality speech llm training and inference.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Low frame-rate speech codec: a codec designed for fast high-quality speech llm training and inference

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:46.705659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:46.705659Z digest=sha256:669a3a4b3bc989f14732ca0086b44926583ec8f278d329f3de70ddf1ca80cef0

Observation 8355f08a-cb85-47f3-8dae-a183f17c899a · outbound

This paper cites A speech-to-speech translation based interface for tourism.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI A speech-to-speech translation based interface for tourism

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:46.827097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:46.827097Z digest=sha256:b5e904ef0e7c54387123893ff58e1e6446a4434677159fdaa60b1e929ec98884

Observation 257d59fa-52f2-480f-942e-00eab9818c01 · outbound

This paper cites Exploring in-context learning of textless speech language model for speech classification tasks.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Exploring in-context learning of textless speech language model for speech classification tasks

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:46.942799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:46.942799Z digest=sha256:743531d365a2afef13a21add19d65a8826647e5019c48529a7cb50fbe8d2111a

Observation ffd1783b-d96b-45b3-875b-cd32e9e3f7cc · outbound

This paper cites Audio Large Language Models Can Be Descriptive Speech Quality Evaluators.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Audio Large Language Models Can Be Descriptive Speech Quality Evaluators

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:47.037457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:47.037457Z digest=sha256:7b1124945812406d835d1266c70be1b3a408c193392bac3aff8dc138e5b61aca

Observation 50a4e804-f797-4395-8160-ed86c96a8184 · outbound

This paper cites Multi-modal generative ai: Multi-modal llm, diffusion and beyond.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Multi-modal generative ai: Multi-modal llm, diffusion and beyond

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:47.131408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:47.131408Z digest=sha256:8d484f240f33aa2b58a2c5ac8d9deb660c0ada9d7bf6de69817c24f595607e73

Observation cfb96779-2c01-4c02-8f85-b926d2500d09 · outbound

This paper cites BLASER: A Text-Free Speech-to-Speech Translation Evaluation Metric.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI BLASER: A Text-Free Speech-to-Speech Translation Evaluation Metric

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:29:09.300688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:28:47.272554Z digest=sha256:df7ca4971d096a6d31963e27105fc89bbd081b2267862a6ba05f2faf91b03271

Observation e0a91960-f798-408d-9327-e419336192c6 · outbound

This paper cites Opportunities and challenges of diffusion models for generative ai.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Opportunities and challenges of diffusion models for generative ai

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:47.395714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:47.395714Z digest=sha256:dc1aca38fea0c1093d45f91635010eacc747d5c12031865f65b06f70cc09199f

Observation d123d686-eb6c-44ea-b2e1-1c3d8cda3ec5 · outbound

This paper cites Speech-to-Speech Translation For A Real-world Unwritten Language.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Speech-to-Speech Translation For A Real-world Unwritten Language

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:29:08.984427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:28:47.506438Z digest=sha256:bb91205362faedbcd3eb4a91cd7afe4c8fbfaab58aea7758f94fad25c6c920f3

Observation 45848d7e-2c1f-42fb-ad7b-5181aebd560e · outbound

This paper cites Beyond Single-Audio: Advancing Multi-Audio Processing in Audio Large Language Models.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Beyond Single-Audio: Advancing Multi-Audio Processing in Audio Large Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:47.620955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:47.620955Z digest=sha256:be8155eb78c358cac40d5c8a232cc117f0bb14eed2709f2234e2bffc4937f508

Observation f85aa4a1-dd53-4ecd-855f-9464586307fb · outbound

This paper cites MAVFlow: Preserving Paralinguistic Elements with Conditional Flow Matching for Zero-Shot AV2AV Multilingual Translation.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI MAVFlow: Preserving Paralinguistic Elements with Conditional Flow Matching for Zero-Shot AV2AV Multilingual Translation

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:29:08.710391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:28:47.738849Z digest=sha256:e375224e0c99da4943b7c6cc31b5b7ad6505d0da97aa6577edde81923f74b66e

Observation 4cd6b87b-9d0a-4756-914e-c97a4f230078 · outbound

This paper cites V2sflow: Video-to-speech generation with speech decomposition and rectified flow.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI V2sflow: Video-to-speech generation with speech decomposition and rectified flow

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:47.864422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:47.864422Z digest=sha256:7ea8eb34e78bc84fb33b13f8aeb1b4bdce4ab239150a3d7a0716a3b5b2d26bbe

Observation ed150bb4-fd36-4975-ba85-8a9f662e12e9 · outbound

This paper cites Qwen2-Audio Technical Report.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Qwen2-Audio Technical Report

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:47.996341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:47.996341Z digest=sha256:2adc6e300602c86374743b4f0d160cfa0af2faec94f44acfdedc9bbd663b4d93

Observation 70e5c870-ac9f-41f7-bfa6-0dd9bddde940 · outbound

This paper cites Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:48.113094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:48.113094Z digest=sha256:7436562ae63e21afa31878de98e2dcad1cd387a6f15022a1ae98595029e6973c

Observation b4b26d0b-0a3d-4474-ab62-54688b122d44 · outbound

This paper cites Diffusion models in vision: A survey.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Diffusion models in vision: A survey

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:48.258509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:48.258509Z digest=sha256:7f54f0788e594c1c5d1b469818009dd104442d417fda25a9f8aaabc0e7566e03

Observation 499a2de5-a6d6-49b1-9bbe-b68e66bb5ecb · outbound

This paper cites Exploring the Benefits of Tokenization of Discrete Acoustic Units.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Exploring the Benefits of Tokenization of Discrete Acoustic Units

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:48.396469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:48.396469Z digest=sha256:145f4ca293791f1b86a941ebc0a700ea831f8501ba63f282341b50ab79ef47a7

Observation aceb2984-3db4-4261-9616-02507bf01c93 · outbound

This paper cites ADIFF: Explaining audio difference using natural language.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI ADIFF: Explaining audio difference using natural language

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:48.510861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:48.510861Z digest=sha256:c90e28b7262ce3dfd75a059750c5061326dce9ecd277326d46cc2da1d6e1d583

Observation ab55b59d-fcab-4bd1-88ca-d70844f0a013 · outbound

This paper cites French- fulfulde textless and cascading speech translation: Towards a dual architecture.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI French- fulfulde textless and cascading speech translation: Towards a dual architecture

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:48.632625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:48.632625Z digest=sha256:bc892dba6ba361accf29607776a996748283d8f21b2b92c5b551c465a6a4fd46

Observation d1013065-7116-412b-8efe-45c6e9a37755 · outbound

This paper cites Textless Speech-to-Speech Translation With Limited Parallel Data.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Textless Speech-to-Speech Translation With Limited Parallel Data

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:29:08.377385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:28:48.766848Z digest=sha256:fc40504cff84e5be0610e5d2f5d918a35e37a163edf9c2f3e3f8a7a1854046c2

Observation a17fd686-12f2-4ec2-b5a1-979ad80e7a4b · outbound

This paper cites PolyVoice: Language Models for Speech to Speech Translation.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI PolyVoice: Language Models for Speech to Speech Translation

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:48.876266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:48.876266Z digest=sha256:16b2e22289bd45dea4d73692c5addac40ee3a213686d566d4322b57ff4703b6b

Observation 0ed51a5e-b3e9-4b28-9310-5c372fc528a5 · outbound

This paper cites LauraGPT: Listen, Attend, Understand, and Regenerate Audio with GPT.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI LauraGPT: Listen, Attend, Understand, and Regenerate Audio with GPT

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:48.957781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:48.957781Z digest=sha256:180147182457b647f4df673407f2d610dfdc6778498bb5ea82bc51ca7d1a5d3a

Observation 4826969a-6b2d-429a-8547-47185ea11b6d · outbound

This paper cites Analyzing Speech Unit Selection for Textless Speech-to-Speech Translation.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Analyzing Speech Unit Selection for Textless Speech-to-Speech Translation

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:29:08.081160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:28:49.074528Z digest=sha256:8f3ed49d42bf3724c1003fda3a9b651554ed7ffb8950cf4034c80509a01bf442

Observation 0eac4a58-04d9-427d-bad9-8e49c446503e · outbound

This paper cites Enhancing expressivity transfer in textless speech-to-speech translation.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Enhancing expressivity transfer in textless speech-to-speech translation

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:49.184864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:49.184864Z digest=sha256:1216a678173bb0303705c045fdbf1ef6928272cb4357cbaafa3f569adbc7e83b

Observation bb15fd2f-4b1b-47a7-8885-81dd702209da · outbound

This paper cites Towards massive parallel corpus creation for hausa-to-english machine translation.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Towards massive parallel corpus creation for hausa-to-english machine translation

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:49.309000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:49.309000Z digest=sha256:5b5bc3388b23a8f8726f34b006dd72a3ccb19e1c22d6ef578b0992e5dcb5beda

Observation 43061353-529b-4bd1-b9e2-4e95189b84e7 · outbound

This paper cites Auditory-visual perception of speech.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Auditory-visual perception of speech

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:49.422521Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:49.422521Z digest=sha256:dd5f31896d9fa6698d088f2e53f6e8d7eea4f2ae7e57fe3feb0cf6d0e4a5dd1d

Observation 91bd5a83-bffa-449d-ab01-199c8b972f30 · outbound

This paper cites Cascade or direct speech translation? a case study.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Cascade or direct speech translation? a case study

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:49.526893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:49.526893Z digest=sha256:4073c8ce67075ee3283040ce08765cc800ecd93e50a658611215ce101109e488

Observation e5b44b4b-c15b-4aa9-b989-426f7342c4ab · outbound

This paper cites CTC-based Non-autoregressive Textless Speech-to-Speech Translation.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI CTC-based Non-autoregressive Textless Speech-to-Speech Translation

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:49.651108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:49.651108Z digest=sha256:6eca2a4996d2576d92382b06347e74a1c4ae822beb73a035126980e728e635fa

Observation b5a56457-cccd-405c-a2f8-3d656195576b · outbound

This paper cites Can We Achieve High-quality Direct Speech-to-Speech Translation without Parallel Speech Data?.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Can We Achieve High-quality Direct Speech-to-Speech Translation without Parallel Speech Data?

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:29:07.737070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:28:49.736451Z digest=sha256:dde203d2107f8c8b5c918f4c295a78605f721cf8d35f3fdec7f65668e3712a42

Observation aef21d98-ff4e-4d1a-b6bc-c81f4bd7ba12 · outbound

This paper cites Generative learning of the solution of parametric partial differential equations using guided diffusion models and virtual observations.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Generative learning of the solution of parametric partial differential equations using guided diffusion models and virtual observations

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:49.823671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:49.823671Z digest=sha256:7e65e7c12247ba020406159a7700bc14643e70aba7dae45ebed3fe259cd0024c

Observation 2623a34d-a06c-4916-b297-75ecf9cc5921 · outbound

This paper cites Unsupervised speech technology for low-resource languages.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Unsupervised speech technology for low-resource languages

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:49.953797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:49.953797Z digest=sha256:f1c91da9574c0dd548546a2fd0f54e4fe111e18ef099e29270fb0c70baadb03d

Observation 0da42731-d665-45ed-a794-9f9946f11ffa · outbound

This paper cites Speech-to-speech translation.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Speech-to-speech translation

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:50.060496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:50.060496Z digest=sha256:36dcbb9f8c8ea68a70b1278d181cd295f2d699c5c5dfe516b87e40caa94b2108

Observation 6e31a667-aa47-4940-9947-2d4e8bbd8fba · outbound

This paper cites Audio Dialogues: Dialogues dataset for audio and music understanding.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Audio Dialogues: Dialogues dataset for audio and music understanding

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:50.157071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:50.157071Z digest=sha256:64a57879fc081d3a5750da3e3ede5b95b3fd8ec0395c8ba284b28276ae79f7fb

Observation 60b14585-6067-444f-a72e-8091a3e33557 · outbound

This paper cites Multilingual Speech-to-Speech Translation into Multiple Target Languages.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Multilingual Speech-to-Speech Translation into Multiple Target Languages

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:50.237516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:50.237516Z digest=sha256:d95d2b5d3035cdb581d8eb9e897f13b43649ff4a6cb0efc7ef0dc2975706a84a

Observation f8cd0254-62a1-47a9-9fda-2f42e927c5f1 · outbound

This paper cites Joint audio and speech understanding.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Joint audio and speech understanding

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:50.328623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:50.328623Z digest=sha256:dc7cb1f5c9d3a19e5e88927d1c9fb14b2e37078b2cac432593d7609b0e96be21

Observation 39fa5909-c21d-46eb-a46f-2ee384003423 · outbound

This paper cites Tibetan–chinese speech-to-speech translation based on dis- crete units.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Tibetan–chinese speech-to-speech translation based on dis- crete units

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:50.423111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:50.423111Z digest=sha256:fd30a57dcfe37609d0d91475730ad667d08090482e4197c223b954ff4ae6c94a

Observation 95e3ae8a-f5a0-4b8e-b1b8-f29255c55b25 · outbound

This paper cites Recent advances in discrete speech tokens: A review.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Recent advances in discrete speech tokens: A review

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:50.514977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:50.514977Z digest=sha256:41de164a4c738a712a772588859754aee732e0d82c35b3e2ab69b096cca5c9b1

Observation 666aa822-1a2a-4348-bae2-6e5fe90247a1 · outbound

This paper cites Direct Speech-to-Speech Neural Machine Translation: A Survey.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Direct Speech-to-Speech Neural Machine Translation: A Survey

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:29:07.434642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:28:50.617509Z digest=sha256:ff34a0642401ea0f9043ebf7882563ec4928a083e92696b47ae6ab695dd39471

Observation 2022b7aa-ae5c-427c-a578-058301bac9b2 · outbound

This paper cites Onellm: One framework to align all modalities with language.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Onellm: One framework to align all modalities with language

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:50.713865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:50.713865Z digest=sha256:971e31e8b5cb2baa9e10012dfe47f4b2511c8eb1d5713abcda8865088b0dd767

Observation fd866ec3-b451-408d-b0a0-8748559e606d · outbound

This paper cites Physics-inspired approaches in generative diffusion models.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Physics-inspired approaches in generative diffusion models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:50.772569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:50.772569Z digest=sha256:233fad68e97318b9f1c1197499cd3cd886ac5552321b9c03f6cc22e06d835bbe

Observation f355817d-0e17-4169-9cbd-8afe0eb53b49 · outbound

This paper cites Exploring In-Context Learning of Textless Speech Language Model for Speech Classification Tasks.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Exploring In-Context Learning of Textless Speech Language Model for Speech Classification Tasks

Reference 48

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:29:07.249632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:28:50.849055Z digest=sha256:3f62fb2ca1888c25bc7e0610c975957a0127215e6a59a07b3322c5ea254b4efd

Observation fbc096ac-803e-475f-9254-687f59be384f · outbound

This paper cites Chain-of-thought prompting for speech translation.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Chain-of-thought prompting for speech translation

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:50.906902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:50.906902Z digest=sha256:b52f9b89cf9b7ca944efd32d252f33ffb82d84b6fb07c4a98ea18e82c5289cd3

Observation a16b2aec-5c45-419e-bdd6-f59e89cb8817 · outbound

This paper cites TranSpeech: Speech-to-Speech Translation With Bilateral Perturbation.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI TranSpeech: Speech-to-Speech Translation With Bilateral Perturbation

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:29:07.050233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:28:50.968083Z digest=sha256:970154c9c7613d80ee7666c5cbb2fdc63cab4d191eeb2bb6ef3031d712bba56b

Observation 645479fc-6505-4d81-8dc5-c3e9f1e7c109 · outbound

This paper cites Textless Acoustic Model with Self-Supervised Distillation for Noise-Robust Expressive Speech-to-Speech Translation.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Textless Acoustic Model with Self-Supervised Distillation for Noise-Robust Expressive Speech-to-Speech Translation

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:29:06.877355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:28:51.039084Z digest=sha256:536893160d45ac3be9e2a6a0436cc63a914e22f992c3b52d05466905ee952a8e

Observation 4f7aa9c0-3d17-439d-b26a-4b34f3d7c969 · outbound

This paper cites Massively multi- lingual forced aligner leveraging self-supervised discrete units.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Massively multi- lingual forced aligner leveraging self-supervised discrete units

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:51.104629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:51.104629Z digest=sha256:fba2d0d5211b230fd38e71b96a8d74a9646515aef8dc7c0899829c5a87eaf20e

Observation 411f20dd-1332-4924-ac1f-3640152ea1a7 · outbound

This paper cites LibriS2S: A German-English Speech-to-Speech Translation Corpus.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI LibriS2S: A German-English Speech-to-Speech Translation Corpus

Reference 53

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:29:06.641946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:28:51.219694Z digest=sha256:4e09e78ed765375fafbef71127879cc53605afea9f7eed956fb3279b648cea71

Observation b3c742b7-aa8e-4ea6-b906-76cf35b79278 · outbound

This paper cites WavChat: A Survey of Spoken Dialogue Models.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI WavChat: A Survey of Spoken Dialogue Models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:51.286122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:51.286122Z digest=sha256:e5689fad4fa6069c36218a4047c3b3bf87e98db7e9b11818b9f9eb557e04d048

Observation 73b92d89-70e8-41c1-936a-f94529374e76 · outbound

This paper cites Can Generative Geospatial Diffusion Models Excel as Discriminative Geospatial Foundation Models?.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Can Generative Geospatial Diffusion Models Excel as Discriminative Geospatial Foundation Models?

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:51.380557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:51.380557Z digest=sha256:48de6018875e4042163c7b5c5899ecc643511e22bc93d4e07caab398fdea4e5c

Observation 2b4fde17-57c6-4b7f-b51f-693b03ae020e · outbound

This paper cites Listra automatic speech translation: English to lingala case study.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Listra automatic speech translation: English to lingala case study

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:51.433996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:51.433996Z digest=sha256:916468e432b0b36dccf75183d04a80e8794fdd1b4b23e8506e25c914d640c1e0

Observation 8164b2a9-f011-4b13-8074-1a8b4d7bb00d · outbound

This paper cites Gdplan: Generative network planning via graph diffusion model.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Gdplan: Generative network planning via graph diffusion model

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:51.582236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:51.582236Z digest=sha256:06e97a6fb3bc4a37663371a5dc79020126ea37f0c30da02352eef56381932504

Observation ffa65f22-3545-4468-a606-d1383c7257be · outbound

This paper cites Direct Punjabi to English speech translation using discrete units.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Direct Punjabi to English speech translation using discrete units

Reference 58

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:29:06.313102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:28:51.726334Z digest=sha256:6f8c890bcc7158b9fd2b46c903e236009a105542fd96e3f2e54cd0251ff466d1

Observation 47a8994c-a3da-440a-93fb-7ce5beed79f2 · outbound

This paper cites Textless unit-to-unit training for many- to-many multilingual speech-to-speech translation.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Textless unit-to-unit training for many- to-many multilingual speech-to-speech translation

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:51.872948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:51.872948Z digest=sha256:4324cd0433f26d4e8fd34aed0ee97cef9ea85693a4dbc63b51d2f04662a32de4

Observation f1df9749-84f9-43fa-a302-e888b5d5135b · outbound

This paper cites Phi dm-dialog: an experimental speech-to-speech dialog translation system.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Phi dm-dialog: an experimental speech-to-speech dialog translation system

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:52.075185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:52.075185Z digest=sha256:33c571be3d3c39c48cfbb5b7ba59e527e3749b265338d0b72764949ffcb7df78

Observation e1bcde44-36e0-4eca-ae94-34ec2c0fdf4f · outbound

This paper cites High-Fidelity Simultaneous Speech-To-Speech Translation.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI High-Fidelity Simultaneous Speech-To-Speech Translation

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:52.222146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:52.222146Z digest=sha256:42c4b5e89628914fd8a53c62107649f735e2ce8ee8885d62c46d758a7f072a64

Observation ee857ae4-18f4-4465-91a5-e7e86832f1ea · outbound

This paper cites Janus-iii: Speech-to-speech translation in multiple languages.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Janus-iii: Speech-to-speech translation in multiple languages

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:52.338122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:52.338122Z digest=sha256:f9c8ed420b4ac0cee95de67f79fc7cb75ca82858a482d50e6ec5a31fc9bb6369

Observation a544d151-2432-4379-afe7-49bd57a39dac · outbound

This paper cites Textless Speech-to-Speech Translation on Real Data.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Textless Speech-to-Speech Translation on Real Data

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:52.430942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:52.430942Z digest=sha256:58f6b4e114ce9570fe83b8af2bc8265b2f1d151db394850dac045301709e04ec

Observation 7bd8c795-7df6-4383-9fdd-0cfa0c5444aa · outbound

This paper cites Video diffusion models are strong video inpainter.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Video diffusion models are strong video inpainter

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:52.488633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:52.488633Z digest=sha256:2e4a9b344ea43b88109972314454011a8be350f701ed1db07c5963bc04b41136

Observation 6f212478-dc0f-4540-8470-5ec1cb2360fd · outbound

This paper cites Speech proportion and accuracy in simultaneous interpretation from english into korean.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Speech proportion and accuracy in simultaneous interpretation from english into korean

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:52.561720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:52.561720Z digest=sha256:c32fb8b3d87529d5d7ea35e08817facb10402e63177f3a6eb1141621a68a69cb

Observation c5462798-f80a-43c7-a9e3-edd0a298c7b6 · outbound

This paper cites Diffusion models for audio restoration: A review [special issue on model-based and data-driven audio signal processing].

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Diffusion models for audio restoration: A review [special issue on model-based and data-driven audio signal processing]

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:52.627257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:52.627257Z digest=sha256:a0269a07d1ecc15c46cdbbc0e4c54f9dbb9cdd0af1416fa23ee628ebbf8e326d

Observation 2bfdd863-1ecb-41df-8902-eb01fd97d21a · outbound

This paper cites Conditional diffusion model for missing value imputation.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Conditional diffusion model for missing value imputation

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:52.698356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:52.698356Z digest=sha256:85c7ce17fd34774324483b1636a6070409239cf0291ca01d59b58c1194c1b0c2

Observation 69f88707-5b02-41bd-b2e2-f5ae64cbdcff · outbound

This paper cites BrainECHO: Semantic Brain Signal Decoding through Vector-Quantized Spectrogram Reconstruction for Whisper-Enhanced Text Generation.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI BrainECHO: Semantic Brain Signal Decoding through Vector-Quantized Spectrogram Reconstruction for Whisper-Enhanced Text Generation

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:52.756373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:52.756373Z digest=sha256:9e766e00a052780025d8ac9d17b1957c508d5446bb7e2fbad75a36cb75fb8680

Observation a861450a-14e7-4ae6-b631-4ebdec81cd59 · outbound

This paper cites Baichuan-Audio: A Unified Framework for End-to-End Speech Interaction.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Baichuan-Audio: A Unified Framework for End-to-End Speech Interaction

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:52.822450Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:52.822450Z digest=sha256:f554c42d45a486bc3a5ee14be1506abb2bc4cc23fedacb57fd76e3d8db42ed87

Observation c7976c58-03ca-45cc-b81f-0ef8ea4a3323 · outbound

This paper cites Textless direct speech-to-speech translation with discrete speech representation.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Textless direct speech-to-speech translation with discrete speech representation

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:52.889256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:52.889256Z digest=sha256:b34a7525f5e64434800812df4d20ddef5dcbaeb7ce53cb5c618308d14d23909c

Observation 396661f3-0e18-4ef1-a595-31a025022a83 · outbound

This paper cites Beyond Words: AuralLLM and SignMST-C for Sign Language Production and Bidirectional Accessibility.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Beyond Words: AuralLLM and SignMST-C for Sign Language Production and Bidirectional Accessibility

Reference 71

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:29:06.009246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:28:52.948623Z digest=sha256:a7e2638f53a8064cf700bb421c3daf6184471dbbcdd851e616c98824d4ef70ae

Observation ccef7ffc-1061-4497-8b00-b6dc666de485 · outbound

This paper cites Align-SLM: Textless Spoken Language Models with Reinforcement Learning from AI Feedback.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Align-SLM: Textless Spoken Language Models with Reinforcement Learning from AI Feedback

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:53.047690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:53.047690Z digest=sha256:fa2f7312ee32df1a3baafdbf1e9f5f2e56afb87c29c723bb0faa1f7feaeeaa4b

Observation 70b4be9b-ed03-4861-b1d3-1846e98fd97e · outbound

This paper cites an unresolved cited work.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Unresolved cited work

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:53.113989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:53.113989Z digest=sha256:34efc408bb9c8a0490219ca0908618def68a119b2978292959570fb71520bf0f

Observation 79af8f6c-b061-4ac3-88c1-3a856ac09d70 · outbound

This paper cites Handdiffuse: generative controllers for two-hand interactions via diffusion models.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Handdiffuse: generative controllers for two-hand interactions via diffusion models

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:53.190421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:53.190421Z digest=sha256:87d9f0b6fd6b8c54b103ceecb698d1718dc6469340fc64d3cfca4c55a62237c1

Observation 0222dac1-65b2-4a42-b3f4-ba5a3b496eda · outbound

This paper cites A Preliminary Exploration with GPT-4o Voice Mode.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI A Preliminary Exploration with GPT-4o Voice Mode

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:53.249841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:53.249841Z digest=sha256:2c76c95aacd43c2b29981d8df329249f1024333ec630618199d1fe92290f836f

Observation 550e152a-e5dc-42bc-8b1f-74d8a62b5a70 · outbound

This paper cites Recent highlights in multilingual and multimodal speech translation.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Recent highlights in multilingual and multimodal speech translation

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:53.332770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:53.332770Z digest=sha256:c2903e091e950cdf977a0c29e93f16c9f1a36e53b8efa3401750fcfbe0211269

Observation a738617c-b7fd-458e-825c-4d7de2be9be0 · outbound

This paper cites Speech-to-speech low-resource translation.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Speech-to-speech low-resource translation

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:53.418689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:53.418689Z digest=sha256:e8606fdd1c6428e1bf8fe16ae3f7570a30b25cb1d1b2e0170bdcda62b6600194

Observation 1b92ec96-557f-43f8-a985-c861c919bd5f · outbound

This paper cites Listening and seeing again: Generative error correction for audio-visual speech recognition.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Listening and seeing again: Generative error correction for audio-visual speech recognition

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:53.480148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:53.480148Z digest=sha256:ad28289c0272dc92ae7ec3b8e89128391c55a51962e0fa8d2007df128b26ae0f

Observation 497f7294-6426-4c88-8ee4-a722d4b124f6 · outbound

This paper cites SLIDE: Integrating Speech Language Model with LLM for Spontaneous Spoken Dialogue Generation.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI SLIDE: Integrating Speech Language Model with LLM for Spontaneous Spoken Dialogue Generation

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:53.542723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:53.542723Z digest=sha256:bf5c999a4ee3d65ce011bbcb83a4606dcd34f348a872ad21b121b71a316d4fca

Observation 51da33f3-43af-4f58-b49d-4360c75d864c · outbound

This paper cites Llamapartialspoof: An llm-driven fake speech dataset simulating disinformation generation.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Llamapartialspoof: An llm-driven fake speech dataset simulating disinformation generation

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:53.615822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:53.615822Z digest=sha256:8d2e9edd2e913d6c434bcf73fbda06c39f24ada53907d925ae2fe555c0eacc86

Observation d8a008de-eb3f-46d5-a025-6441456e2a81 · outbound

This paper cites Build llm-based zero-shot streaming tts system with cosyvoice.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Build llm-based zero-shot streaming tts system with cosyvoice

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:53.691843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:53.691843Z digest=sha256:2d0ff86e0f16fe672b7ef82fccef788ac79edcd88783e5e6cf82b8abbe802e1b

Observation 7da2b580-0764-4db3-813a-999550cfc9a4 · outbound

This paper cites Auto-avsr: Audio-visual speech recognition with automatic labels.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Auto-avsr: Audio-visual speech recognition with automatic labels

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:53.754082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:53.754082Z digest=sha256:9b31c443d3349bdc0b2d5e98f9bc40e5a97cb75a999b64129a55cae4be547647

Observation 7fa3d4a9-2460-499f-a34e-4f104d4c58c9 · outbound

This paper cites Real-Time Textless Dialogue Generation.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Real-Time Textless Dialogue Generation

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:53.813975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:53.813975Z digest=sha256:2c1bb6292d41efc7c7230034c0ec115dd70933b05da4c4559fe582aa18d11916

Observation 3cb01864-40b8-4d19-ae41-18a246831544 · outbound

This paper cites Slamming: Training a Speech Language Model on One GPU in a Day.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Slamming: Training a Speech Language Model on One GPU in a Day

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:53.898225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:53.898225Z digest=sha256:a5f9e4e141dfe5da096d0d2756bf08ec116cc997ae46df3a0fa349cc171c9b96

Observation 809a6685-b873-4b14-9281-a9ece5f55ba6 · outbound

This paper cites Enhancing Low-Resource Language and Instruction Following Capabilities of Audio Language Models.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Enhancing Low-Resource Language and Instruction Following Capabilities of Audio Language Models

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:53.960789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:53.960789Z digest=sha256:3d83de2e93107915aae18f28a87970a70786c2c93db3d264e49884e4e889d0f2

Observation f53f924c-240e-4287-b891-99000f3eec12 · outbound

This paper cites Make some noise: Towards llm audio reasoning and generation using sound tokens.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Make some noise: Towards llm audio reasoning and generation using sound tokens

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:54.025981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:54.025981Z digest=sha256:9b7bc6ed8a3dfee182b1e3cf9d2f4f1a5c0413ad2f298b229e46077df1f51fbd

Observation 8e180aeb-c0ca-4e72-9646-087233593a67 · outbound

This paper cites Deep networks as denoising algorithms: Sample-efficient learning of diffusion models in high-dimensional graphical models.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Deep networks as denoising algorithms: Sample-efficient learning of diffusion models in high-dimensional graphical models

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:54.074608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:54.074608Z digest=sha256:7b94308115601121ade0caf89517f4dc47d940450e589e4cbc550686bd27f654

Observation e4e61795-1735-49f3-b201-3cbf5270002a · outbound

This paper cites Amharic speech recognition for speech translation.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Amharic speech recognition for speech translation

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:54.129617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:54.129617Z digest=sha256:fd69dd4eb508ef39c3ab3cea87233aaa5f57dceca662c1aacf0babceee9307c0

Observation 85a5e050-475a-4d8f-a090-e283bb7612b0 · outbound

This paper cites Parrot: Autoregressive spoken dialogue language modeling with decoder-only transform- ers.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Parrot: Autoregressive spoken dialogue language modeling with decoder-only transform- ers

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:54.215355Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:54.215355Z digest=sha256:e295935adc7e735c8071803a7aedba6449d9fb7b8682044abf98e47d8ce685d7

Observation beb0db37-1d76-4902-a215-7357bd88f8eb · outbound

This paper cites Towards to a direct speech to speech for endangered languages in africa.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Towards to a direct speech to speech for endangered languages in africa

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:54.281049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:54.281049Z digest=sha256:f4c82573b65d42c32e54485a6af086dae8cf67095e76800b91e35204f1a79297

Observation a87d303b-cef5-4093-a2d6-25f6d4b7597b · outbound

This paper cites A Unit-based System and Dataset for Expressive Direct Speech-to-Speech Translation.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI A Unit-based System and Dataset for Expressive Direct Speech-to-Speech Translation

Reference 91

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:29:05.757265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:28:54.358626Z digest=sha256:133e75d281e45f65cd8e3842693ef227f2985fe9c51559e6d751aacab88f2a3f

Observation 5d30e757-ba22-422a-b20a-df145a1b26f0 · outbound

This paper cites Spoken Question Answering and Speech Continuation Using Spectrogram-Powered LLM.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Spoken Question Answering and Speech Continuation Using Spectrogram-Powered LLM

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:54.440358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:54.440358Z digest=sha256:7aa727194a0e4b73215d30ffd98bfd6ead5a6da47554cd41187dfb4533d30d4f

Observation 404457b9-4d9e-466b-89c6-ef2fde59496c · outbound

This paper cites Towards real-time multilingual multimodal speech-to-speech translation.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Towards real-time multilingual multimodal speech-to-speech translation

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:54.514589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:54.514589Z digest=sha256:7f952a8a27379dc61eaf5e7a3506f301281b68d79f4d6c9cb311a0b570d97967

Observation 29d563db-e77d-4365-8701-2432871951b3 · outbound

This paper cites One Model, Many Languages: Meta-learning for Multilingual Text-to-Speech.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI One Model, Many Languages: Meta-learning for Multilingual Text-to-Speech

Reference 94

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:29:05.595083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:28:54.554883Z digest=sha256:a3274b0217d290eae797f4c3508c7e45d1f1dd45e7f1a15eaae0af8412a8abb1

Observation 6f82a242-5629-4085-b9f6-6be3c3184587 · outbound

This paper cites Spoken Language Modeling from Raw Audio.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Spoken Language Modeling from Raw Audio

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:54.618527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:54.618527Z digest=sha256:0d0876ba567e33115028217c2ee945bfe103dcfef2915be7ebf7778cbd960ca0

Observation a296a4c6-ff87-4be2-ac85-65a5381ba322 · outbound

This paper cites Visually Grounded Speech Models for Low-resource Languages and Cognitive Modelling.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Visually Grounded Speech Models for Low-resource Languages and Cognitive Modelling

Reference 96

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:29:05.415512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:28:54.661961Z digest=sha256:bf7a60368a899067a5bfe7cdfc4512928a28573fe03e54892062af65ef32c5a7

Observation 04064344-4acb-46fc-8596-1a702b74407a · outbound

This paper cites Verbmobil: The use of prosody in the linguistic components of a speech understanding system.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Verbmobil: The use of prosody in the linguistic components of a speech understanding system

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:54.778940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:54.778940Z digest=sha256:1c6d00f703179b122e393f4f31ab477621e2b61ce02dce6958180a02833559aa

Observation 34d099fb-c96f-4cec-85e7-4e0ec50d7b47 · outbound

This paper cites Phonology-Guided Speech-to-Speech Translation for African Languages.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Phonology-Guided Speech-to-Speech Translation for African Languages

Reference 98

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:29:05.278205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:28:54.842017Z digest=sha256:ce34b2420f920150a655e16d437e702e644bdbdc94644553a5add5005607af70

Observation 899aff8c-87f0-48a0-b348-c1531fdc2af2 · outbound

This paper cites Let's Go Real Talk: Spoken Dialogue Model for Face-to-Face Conversation.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Let's Go Real Talk: Spoken Dialogue Model for Face-to-Face Conversation

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:54.880977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:54.880977Z digest=sha256:14c118bf03b44f160537f8d61908fb3f66429251925f89d258a6f93babaf61ea

Observation 69f5a34b-30cf-4b75-8c5b-9544d342ba84 · outbound

This paper cites Long-Form Speech Generation with Spoken Language Models.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Long-Form Speech Generation with Spoken Language Models

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:54.976320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:54.976320Z digest=sha256:1106e888e03da68db6e57f17f54d0abfbc997ad304ad299f2f914bc1c52753c6

Pith citing papers

Observation b97f05c3-f9ac-40d9-8c1e-c8c2a4cb0a72 · inbound

Achieving Generational Peace in Mali through Intergenerational Mean-Field-Type Game-based Incentives cites this paper.

Achieving Generational Peace in Mali through Intergenerational Mean-Field-Type Game-based Incentives Breaking the Barriers of Text-Hungry and Audio-Deficient AI

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-21T00:59:19.261302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-21T00:58:13.403170Z digest=sha256:d6ecb0944e20577f1f551af96e2d5104a1862d9b9e838b6c2559b96cf1fcdd85

Observation e83bdcdb-5e9f-4667-9220-e8c24470eaa5 · inbound

Mitigating Polycentric Conflict-Trap Risk in Mali via Intergenerational Volterra Mean-Field-Type Games cites this paper.

Mitigating Polycentric Conflict-Trap Risk in Mali via Intergenerational Volterra Mean-Field-Type Games Breaking the Barriers of Text-Hungry and Audio-Deficient AI

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-07-04T09:39:46.381030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T09:46:15.701501Z digest=sha256:9dfb13de92841b2a4890905ac6ba88d61c3d75e333dcfa157a89bcfab2d2c4f1

Observation f7b6c78e-a369-4e67-867a-2a4f67fa9319 · inbound

Risk-Aware Information Theory cites this paper.

Risk-Aware Information Theory Breaking the Barriers of Text-Hungry and Audio-Deficient AI

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-07-04T09:39:46.873728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T09:41:41.049881Z digest=sha256:2ab651b7889ffabf1c08ab8ad656c73421370542b715961427ce002f95d4e981