Pith. sign in

Paper Citation Record · LEDGER

LauraGPT: Listen, Attend, Understand, and Regenerate Audio with GPT

As of 19 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 38 inbound Pith citation observations for arXiv:2310.04673.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2310.04673 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 38 of 38 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 38 of 38 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T06:05:11.370627Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T12:19:49.665222Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 1bbd4f02-3816-436a-89c4-61a94360561b · inbound

DASB - Discrete Audio and Speech Benchmark cites this paper.

DASB - Discrete Audio and Speech Benchmark LauraGPT: Listen, Attend, Understand, and Regenerate Audio with GPT

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-24T00:28:39.499589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-24T00:26:57.419537Z digest=sha256:b77d743d94ed3374dce34230bddb471f49d3c577e09ffa2a0182cc59b678c29c

Observation 6ff3ebd0-44db-421a-9894-2979d38956c8 · inbound

A Comparative Study of Discrete Speech Tokens for Semantic-Related Tasks with Large Language Models cites this paper.

A Comparative Study of Discrete Speech Tokens for Semantic-Related Tasks with Large Language Models LauraGPT: Listen, Attend, Understand, and Regenerate Audio with GPT

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:18.682101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:18.682101Z digest=sha256:57cd4afccab63b509ca88c232bed37ca799a9feea74e46b59ac1834f8a10bd1f

Observation 5a316dd0-6089-4e12-b772-fe85c4600f4e · inbound

WavChat: A Survey of Spoken Dialogue Models cites this paper.

WavChat: A Survey of Spoken Dialogue Models LauraGPT: Listen, Attend, Understand, and Regenerate Audio with GPT

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-12T20:13:57.204938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:13:57.204938Z digest=sha256:face3f1b8e5825e835f6cca4c5b7269468742bffe1b2c9818f9d04ba8c3ab7e3

Observation fedfc8f9-0c10-4bc1-80ab-eae0ffdfc593 · inbound

AlignFormer: Modality Matching Can Achieve Better Zero-shot Instruction-Following Speech-LLM cites this paper.

AlignFormer: Modality Matching Can Achieve Better Zero-shot Instruction-Following Speech-LLM LauraGPT: Listen, Attend, Understand, and Regenerate Audio with GPT

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T04:44:25.838344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:44:25.838344Z digest=sha256:aacd689545559729de8d5276668e8c69f5b78e9a9a71bef7583effc4d0f7cd89

Observation 4174d551-b364-4fdd-b9c9-6ac3e7c1a9f5 · inbound

Comprehensive Audio Query Handling System with Integrated Expert Models and Contextual Understanding cites this paper.

Comprehensive Audio Query Handling System with Integrated Expert Models and Contextual Understanding LauraGPT: Listen, Attend, Understand, and Regenerate Audio with GPT

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T21:56:41.844752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T21:56:41.844752Z digest=sha256:2dd3f98d57be898ffe059af3ad1f5eab3d2c9d6ef1a58ef8e5eca5526767e5ec

Observation ee79a058-2e8c-4acf-9ec8-70da2858fa0a · inbound

TouchTTS: An Embarrassingly Simple TTS Framework that Everyone Can Touch cites this paper.

TouchTTS: An Embarrassingly Simple TTS Framework that Everyone Can Touch LauraGPT: Listen, Attend, Understand, and Regenerate Audio with GPT

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T18:08:22.096260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:08:22.096260Z digest=sha256:07ada1087ae9e7c164e54f60b368cd67c8aa7c7cef443b55041adb1672049a2d

Observation 2bc23168-9306-4c21-9b9c-417df8767ec7 · inbound

YingSound: Video-Guided Sound Effects Generation with Multi-modal Chain-of-Thought Controls cites this paper.

YingSound: Video-Guided Sound Effects Generation with Multi-modal Chain-of-Thought Controls LauraGPT: Listen, Attend, Understand, and Regenerate Audio with GPT

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T17:17:49.025672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:17:49.025672Z digest=sha256:ca4f446300b3fad12616933c566b10494ff1f9c9fdd733fa90270ddb87a8931e

Observation aa87ad16-16f0-4646-883c-dcc999720988 · inbound

SilVar: Speech Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization cites this paper.

SilVar: Speech Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization LauraGPT: Listen, Attend, Understand, and Regenerate Audio with GPT

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T10:18:36.014991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:18:36.014991Z digest=sha256:d8df67803cc7669dac100bb6332eb97bfaea07e5e87f84084ea02479087f9b4c

Observation 5a9409a4-187f-48ff-9a6d-3cffe028e505 · inbound

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey cites this paper.

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey LauraGPT: Listen, Attend, Understand, and Regenerate Audio with GPT

Reference 102

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:01.581000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:01.581000Z digest=sha256:68e9b278a292c8d54aa0af8825cbf9b3821919300537093ec578128f7f1a842c

Observation 918f8e07-742d-4f2e-902c-5b6441582155 · inbound

CodecFake+: Codec-Based Resynthesized Data as a Proxy for Detecting CodecFake Speech cites this paper.

CodecFake+: Codec-Based Resynthesized Data as a Proxy for Detecting CodecFake Speech LauraGPT: Listen, Attend, Understand, and Regenerate Audio with GPT

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-10T20:34:37.606605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:34:37.606605Z digest=sha256:5a6f29238a4950183e223fe6038dbbbb6c82e7cab34910411bec4766503876f9

Observation ce8b6928-7725-4f9a-ad72-f4fcab158215 · inbound

A Non-autoregressive Model for Joint STT and TTS cites this paper.

A Non-autoregressive Model for Joint STT and TTS LauraGPT: Listen, Attend, Understand, and Regenerate Audio with GPT

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T20:15:28.356241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:15:28.356241Z digest=sha256:a5da55b14da6b23d8c8335917e1b43db5e2e609573147c2489fb07f3744c050e

Observation 9377f019-0df1-4fc9-84dd-54616c730186 · inbound

Muyan-TTS: A Trainable Text-to-Speech Model Optimized for Podcast Scenarios with a $50K Budget cites this paper.

Muyan-TTS: A Trainable Text-to-Speech Model Optimized for Podcast Scenarios with a $50K Budget LauraGPT: Listen, Attend, Understand, and Regenerate Audio with GPT

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-16T06:05:11.370627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T06:05:11.370627Z digest=sha256:f4fa46bb9c87986950969d444bd04dacf3ca46bf6459534bc21936ede14fc685

Observation f5e1c955-edc6-462f-bfaf-0e48780e59b4 · inbound

Contextual Paralinguistic Data Creation for Multi-Modal Speech-LLM: Data Condensation and Spoken QA Generation cites this paper.

Contextual Paralinguistic Data Creation for Multi-Modal Speech-LLM: Data Condensation and Spoken QA Generation LauraGPT: Listen, Attend, Understand, and Regenerate Audio with GPT

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T20:17:33.399189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:17:33.399189Z digest=sha256:b606d49b2b88c3e03e8f8b56ef8f40ce4bd66e85ee298356727361d9f1968bbe

Observation 454251e3-9fde-4541-94bc-08c0f9dd0653 · inbound

Improving Noise Robustness of LLM-based Zero-shot TTS via Discrete Acoustic Token Denoising cites this paper.

Improving Noise Robustness of LLM-based Zero-shot TTS via Discrete Acoustic Token Denoising LauraGPT: Listen, Attend, Understand, and Regenerate Audio with GPT

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:38.831798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:13:38.831798Z digest=sha256:26f7f51296da455afb5f845340b862f5cba6e363880c5c49ef260c0c320930f1

Observation 6afe8582-3ffc-4312-992a-edafba17ca50 · inbound

Towards Reliable Large Audio Language Model cites this paper.

Towards Reliable Large Audio Language Model LauraGPT: Listen, Attend, Understand, and Regenerate Audio with GPT

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T14:23:45.671049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:23:45.671049Z digest=sha256:f5fad0b8f3f85e836f594afa3ab838193346c4ec0d7eb452c8c3fb183a1a5eba

Observation 6279db10-5903-4d71-adf5-822e051483e3 · inbound

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation cites this paper.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation LauraGPT: Listen, Attend, Understand, and Regenerate Audio with GPT

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:53.892566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:53.892566Z digest=sha256:020a6a78c23268ab03f5ec4e79f0a9cfd1bf4df9d760ff4d99d440ceb89f8a9a

Observation 5ffb8d62-ec18-4574-8387-0ea3b3aaf67b · inbound

Leveraging LLM for Stuttering Speech: A Unified Architecture Bridging Recognition and Event Detection cites this paper.

Leveraging LLM for Stuttering Speech: A Unified Architecture Bridging Recognition and Event Detection LauraGPT: Listen, Attend, Understand, and Regenerate Audio with GPT

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T13:21:55.024547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:21:55.024547Z digest=sha256:66f02ee7c0dc77856a99bf8ac737aa213d2df800c1de18a6fd1e10d738714d73

Observation 0ed51a5e-b3e9-4b28-9310-5c372fc528a5 · inbound

Breaking the Barriers of Text-Hungry and Audio-Deficient AI cites this paper.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI LauraGPT: Listen, Attend, Understand, and Regenerate Audio with GPT

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:48.957781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:48.957781Z digest=sha256:afccdbc3790fb8dbafc8d12b5eae62827cac5f313c1a9c6272ec4c8d60717382

Observation 76296d15-6b1d-4f77-a815-ea0fe44a5d52 · inbound

A Variational Framework for Improving Naturalness in Generative Spoken Language Models cites this paper.

A Variational Framework for Improving Naturalness in Generative Spoken Language Models LauraGPT: Listen, Attend, Understand, and Regenerate Audio with GPT

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T00:15:44.732151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:15:44.732151Z digest=sha256:929ad5723028bb03a639aec79bf0723fa59787c185ac1991abf9f9679c2e53f1

Observation 1dc0d152-f8f7-4e86-bc7e-a8fbaa3b0b19 · inbound

BoSS: Beyond-Semantic Speech cites this paper.

BoSS: Beyond-Semantic Speech LauraGPT: Listen, Attend, Understand, and Regenerate Audio with GPT

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T14:50:04.247739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:50:04.247739Z digest=sha256:7a660912916cee1c5b8f3d535832e8bfc00c0a9231db99071f2727ddbc26bdd8

Observation bafe0a09-d20e-494d-8a8f-e12641a163f3 · inbound

Self-Improvement for Audio Large Language Model using Unlabeled Speech cites this paper.

Self-Improvement for Audio Large Language Model using Unlabeled Speech LauraGPT: Listen, Attend, Understand, and Regenerate Audio with GPT

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T17:52:49.822351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:52:49.822351Z digest=sha256:0ddd3f612bd355ed5d049c0e3f38a9a9a11f465fa527d84613f70cf161365d62

Observation e60264e3-2f83-4ec8-8afa-5a7ab94d894e · inbound

MECAT: A Multi-Experts Constructed Benchmark for Fine-Grained Audio Understanding Tasks cites this paper.

MECAT: A Multi-Experts Constructed Benchmark for Fine-Grained Audio Understanding Tasks LauraGPT: Listen, Attend, Understand, and Regenerate Audio with GPT

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-19T02:41:59.662863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-19T02:41:52.996457Z digest=sha256:d3dd8aad07ce91a16cb20123b5558fb2674f5495cc07e0304ce231c8e34db342

Observation f7bf4d08-aa16-43e5-963e-0f2bc451e6a3 · inbound

LLaSO: A Foundational Framework for Reproducible Research in Large Language and Speech Model cites this paper.

LLaSO: A Foundational Framework for Reproducible Research in Large Language and Speech Model LauraGPT: Listen, Attend, Understand, and Regenerate Audio with GPT

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-05T17:56:51.682540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T17:56:51.682540Z digest=sha256:99a42bf1130500a05955a1fd2818429c256d4c6d96e519e980f867b1c2e8f4b6

Observation 8c6a3ded-de9b-4d0d-8255-efc9e2389adb · inbound

Speech Discrete Tokens or Continuous Features? A Comparative Analysis for Spoken Language Understanding in SpeechLLMs cites this paper.

Speech Discrete Tokens or Continuous Features? A Comparative Analysis for Spoken Language Understanding in SpeechLLMs LauraGPT: Listen, Attend, Understand, and Regenerate Audio with GPT

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T16:46:47.765382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:46:47.765382Z digest=sha256:0c0bb44a0ccfdb5bedaa48029a22558febbb906cdb8aa996457df7d1469aec64

Observation f7c8791c-cb2f-4c63-b8b0-0cad294cf213 · inbound

Group Relative Policy Optimization for Speech Recognition cites this paper.

Group Relative Policy Optimization for Speech Recognition LauraGPT: Listen, Attend, Understand, and Regenerate Audio with GPT

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T12:07:21.313547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:07:21.313547Z digest=sha256:0cd6fdacc0a030c4207dbf416be48e141808b378cc152e1c3baba7f3c2926d02

Observation 4ac6042f-557b-4108-91ec-10fe944b8558 · inbound

FireRedChat: A Pluggable, Full-Duplex Voice Interaction System with Cascaded and Semi-Cascaded Implementations cites this paper.

FireRedChat: A Pluggable, Full-Duplex Voice Interaction System with Cascaded and Semi-Cascaded Implementations LauraGPT: Listen, Attend, Understand, and Regenerate Audio with GPT

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-04T23:33:04.346067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:33:04.346067Z digest=sha256:a7ae411df0b8e369e0e7ac0724b330a88b5f1ce693a02fa093ade9a557e75c66

Observation 3ed13a00-3143-4aea-8a16-0418fe098b38 · inbound

CodecSep: Prompt-Driven Universal Sound Separation on Neural Audio Codec Latents cites this paper.

CodecSep: Prompt-Driven Universal Sound Separation on Neural Audio Codec Latents LauraGPT: Listen, Attend, Understand, and Regenerate Audio with GPT

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-18T16:46:37.612641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-18T16:44:32.104114Z digest=sha256:c99abc1f13ed76867e80c0c0faf0c6c600af9b13fbf40009d9417224b0a21957

Observation ccb6452f-930e-411b-940d-9326f1634eef · inbound

CodecSep: Prompt-Driven Universal Sound Separation on Neural Audio Codec Latents cites this paper.

CodecSep: Prompt-Driven Universal Sound Separation on Neural Audio Codec Latents LauraGPT: Listen, Attend, Understand, and Regenerate Audio with GPT

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T16:48:31.671061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T16:48:31.671061Z digest=sha256:771ec65549d3fac15e9b8065c971a5c1b03a19116b03ea0def1e8cc6b56db357

Observation 16ccd22c-94b2-4b0e-b467-4436db4b9cf0 · inbound

UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models cites this paper.

UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models LauraGPT: Listen, Attend, Understand, and Regenerate Audio with GPT

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T11:29:31.571249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T11:29:31.571249Z digest=sha256:28b6c23892a3599e6843ccaf9ef613b44f750ac0cdd287b9b18e143348ff9fea

Observation 3bee28f1-6bc6-4620-9ba0-da05ee3d0f15 · inbound

Discriminative-Generative Target Speaker Extraction with Decoder-Only Language Models cites this paper.

Discriminative-Generative Target Speaker Extraction with Decoder-Only Language Models LauraGPT: Listen, Attend, Understand, and Regenerate Audio with GPT

Reference 77

Resolution
verified exact
arxiv_id, observed 2026-05-21T15:44:14.736674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-21T15:42:57.766341Z digest=sha256:defc98461102fb3014354fb6d85b3625aebfb4419df74db0fc194c22ba9fee15

Observation 4c30ce28-6c63-4bcb-afdb-7e8437470196 · inbound

StarTSE: Towards Streaming Target Speaker Extraction via Chunk-wise Interleaved Splicing of Autoregressive Language Model cites this paper.

StarTSE: Towards Streaming Target Speaker Extraction via Chunk-wise Interleaved Splicing of Autoregressive Language Model LauraGPT: Listen, Attend, Understand, and Regenerate Audio with GPT

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-10T01:04:50.361710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-10T00:59:51.082759Z digest=sha256:38d5fa42fe288b9711525112ada28291c86b6a7d0b678e973f7728fb4cdbbf52

Observation 0b217a86-fdfc-4ba8-9606-2efad0dde764 · inbound

PairAlign: A Framework for Sequence Tokenization via Self-Alignment with Applications to Audio Tokenization cites this paper.

PairAlign: A Framework for Sequence Tokenization via Self-Alignment with Applications to Audio Tokenization LauraGPT: Listen, Attend, Understand, and Regenerate Audio with GPT

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:16:08.571872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-08T12:25:52.847432Z digest=sha256:ff04d386b69d5ba11d380dbb97a4c9837393c8ce5bcca7ebc1d3657aa353260d

Observation d48d1aa7-664a-4398-a2e2-60701b5e7f1d · inbound

PairAlign: A Framework for Sequence Tokenization via Self-Alignment with Applications to Audio Tokenization cites this paper.

PairAlign: A Framework for Sequence Tokenization via Self-Alignment with Applications to Audio Tokenization LauraGPT: Listen, Attend, Understand, and Regenerate Audio with GPT

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-06-30T23:15:07.781578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-30T23:14:32.494076Z digest=sha256:3227594cc8a4c14f244914050c6910e76dca56af635ae1716271b5a016aa7d08

Observation 645583c1-f047-49f7-a436-2bf7b10d3686 · inbound

A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook cites this paper.

A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook LauraGPT: Listen, Attend, Understand, and Regenerate Audio with GPT

Reference 104

Resolution
verified exact
arxiv_id, observed 2026-05-21T07:39:48.950772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-21T07:38:23.099479Z digest=sha256:1d0e599dfa9b5edeac42fd7bcd2b987c856b57027a70fbfd6f0f1dff0e5ed881

Observation 85f3c95e-ff73-4f5e-84ec-0de7f1e6f0c3 · inbound

From Prompts to Context: An Ontology-Driven Framework for Human-Generative AI Collaboration cites this paper.

From Prompts to Context: An Ontology-Driven Framework for Human-Generative AI Collaboration LauraGPT: Listen, Attend, Understand, and Regenerate Audio with GPT

Reference 37

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T05:53:08.884552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-29T05:49:07.096866Z digest=sha256:9d920c3e763ce7604ef6b1520885df473df00f8763775503adf68714a540dd21

Observation 99f2961a-f65e-42c2-aab2-68a017066e01 · inbound

IRAF: Interference-Resilient Adaptive Fusion for Noise-Robust End-to-End Full-Duplex Spoken Dialogue Systems cites this paper.

IRAF: Interference-Resilient Adaptive Fusion for Noise-Robust End-to-End Full-Duplex Spoken Dialogue Systems LauraGPT: Listen, Attend, Understand, and Regenerate Audio with GPT

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-07-02T15:37:06.822469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-27T23:42:38.203116Z digest=sha256:ad72b3f9df871eed632adae1aade1142c212383aee854d7fb86274ffb6caebaf

Observation 744ba060-03ab-4eb7-9ab3-ecaf55f849e7 · inbound

FlowTTS-GRPO: Online Reinforcement Learning with Multi-Objective Reward Optimization for Flow-Matching Based Text-to-Speech cites this paper.

FlowTTS-GRPO: Online Reinforcement Learning with Multi-Objective Reward Optimization for Flow-Matching Based Text-to-Speech LauraGPT: Listen, Attend, Understand, and Regenerate Audio with GPT

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-07-04T12:19:49.666647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-26T07:02:36.499424Z digest=sha256:3467c9944d9d60652c05d22b3437a4d1f0f2baef3c695350b27d189c6b7e8b47

Observation 7256d16a-5279-4882-a7d4-39165b006d82 · inbound

FlowTTS-GRPO: Online Reinforcement Learning with Multi-Objective Reward Optimization for Flow-Matching Based Text-to-Speech cites this paper.

FlowTTS-GRPO: Online Reinforcement Learning with Multi-Objective Reward Optimization for Flow-Matching Based Text-to-Speech LauraGPT: Listen, Attend, Understand, and Regenerate Audio with GPT

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-12T12:44:20.831164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T12:44:20.831164Z digest=sha256:bc398f682a6ecba5867d864d31a3578501038eec082a12d987761f272ca46c62