Pith. sign in

Paper Citation Record · LEDGER

LLaMA-Omni2: LLM-based Real-time Spoken Chatbot with Autoregressive Streaming Speech Synthesis

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 14 inbound Pith citation observations for arXiv:2505.02625.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.02625 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 14 of 14 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 14 of 14 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:33:19.577773Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

1
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation f5e67540-2a28-4247-9b5c-61fabd5ceea5 · inbound

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model cites this paper.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model LLaMA-Omni2: LLM-based Real-time Spoken Chatbot with Autoregressive Streaming Speech Synthesis

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:19.577773Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:33:19.577773Z digest=sha256:c34c34c181168f9930f6cee1bc94d767196b8ac3496f9d1c7a5f05c6ba6fe124

Observation dc1e401d-6272-4516-afe5-6481c35deecb · inbound

ChipChat: Low-Latency Cascaded Conversational Agent in MLX cites this paper.

ChipChat: Low-Latency Cascaded Conversational Agent in MLX LLaMA-Omni2: LLM-based Real-time Spoken Chatbot with Autoregressive Streaming Speech Synthesis

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T15:52:43.998471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:52:43.998471Z digest=sha256:4f66400e7182d8bcbff969c5ee21a0e71d3037f3c6fba07fa49e856b6900fff0

Observation 1497d69c-0b13-4a0c-8c71-a5d4043ee17c · inbound

StableToken: A Noise-Robust Semantic Speech Tokenizer for Resilient SpeechLLMs cites this paper.

StableToken: A Noise-Robust Semantic Speech Tokenizer for Resilient SpeechLLMs LLaMA-Omni2: LLM-based Real-time Spoken Chatbot with Autoregressive Streaming Speech Synthesis

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-18T13:01:24.365147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-18T12:57:04.450462Z digest=sha256:2ec5bca0fef298c9227839cd866479df5ca8e67982e1e31a8ef5337425780b3b

Observation d503e0da-a3f1-44fe-a927-eed086ae7943 · inbound

Mind-Paced Speaking: A Dual-Brain Approach to Real-Time Reasoning in Spoken Language Models cites this paper.

Mind-Paced Speaking: A Dual-Brain Approach to Real-Time Reasoning in Spoken Language Models LLaMA-Omni2: LLM-based Real-time Spoken Chatbot with Autoregressive Streaming Speech Synthesis

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-18T07:46:03.589866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-18T07:43:23.913399Z digest=sha256:b31cd7c49ebf0ec0febce0338feafeb7210cf12093074b6c80bc9859dd615fc6

Observation 0b2bf6c6-73b0-46cb-9d81-f1357ac89c8f · inbound

Maximizing Local Entropy Where It Matters: Prefix-Aware Localized LLM Unlearning cites this paper.

Maximizing Local Entropy Where It Matters: Prefix-Aware Localized LLM Unlearning LLaMA-Omni2: LLM-based Real-time Spoken Chatbot with Autoregressive Streaming Speech Synthesis

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T12:25:01.710490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T12:25:01.710490Z digest=sha256:87269f95ff333a8e268bb640b2d1b6ad7ca748f63bdb4ec6c5d5347e4df8d89b

Observation 2c54e787-bd4b-49e5-8bd5-f205b4baa4a4 · inbound

Character Beyond Speech: Leveraging Role-Playing Evaluation in Audio Large Language Models via Reinforcement Learning cites this paper.

Character Beyond Speech: Leveraging Role-Playing Evaluation in Audio Large Language Models via Reinforcement Learning LLaMA-Omni2: LLM-based Real-time Spoken Chatbot with Autoregressive Streaming Speech Synthesis

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-10T14:25:30.047240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T14:22:25.660785Z digest=sha256:3749cd24cf7abc13d9f49d5565ba0bba329962403db3c1bf432a8211ecb926aa

Observation 5ff634e1-5961-4926-8e3d-76138e3aa974 · inbound

HalluAudio: A Comprehensive Benchmark for Hallucination Detection in Large Audio-Language Models cites this paper.

HalluAudio: A Comprehensive Benchmark for Hallucination Detection in Large Audio-Language Models LLaMA-Omni2: LLM-based Real-time Spoken Chatbot with Autoregressive Streaming Speech Synthesis

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T13:31:03.386871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T01:34:54.375266Z digest=sha256:2c2c660682cf76c93a6c62d554eff36f2941547110e2ec5fa974cd0fef1ff0ac

Observation 6c489d98-1bb1-4ec1-8352-cbfaa2e5f95b · inbound

SpeechParaling-Bench: A Comprehensive Benchmark for Paralinguistic-Aware Speech Generation cites this paper.

SpeechParaling-Bench: A Comprehensive Benchmark for Paralinguistic-Aware Speech Generation LLaMA-Omni2: LLM-based Real-time Spoken Chatbot with Autoregressive Streaming Speech Synthesis

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-10T00:39:48.480242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T00:39:04.303837Z digest=sha256:d6f2f1c2c6fb4584615a09a8b1421ed3bacf4f2342b76595caa981c5193012ad

Observation 9d99baf3-d54d-4dcc-af2b-3f8950b489ac · inbound

PolySpeech-100: A Large-Scale Benchmark for Speech Understanding Across 100+ Languages and Dialects cites this paper.

PolySpeech-100: A Large-Scale Benchmark for Speech Understanding Across 100+ Languages and Dialects LLaMA-Omni2: LLM-based Real-time Spoken Chatbot with Autoregressive Streaming Speech Synthesis

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-06-28T17:52:26.993555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T17:44:07.669223Z digest=sha256:7a3415437a7214806bfd000483cefdd743bece8c969e4d44b13cbb783aba2252

Observation 023d51a6-739e-4ff4-b00a-dabae30b93a1 · inbound

Audio Interaction Model cites this paper.

Audio Interaction Model LLaMA-Omni2: LLM-based Real-time Spoken Chatbot with Autoregressive Streaming Speech Synthesis

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-02T10:46:52.382243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T04:57:05.062465Z digest=sha256:a2931a1253a8df5f605416e6f35bedb2ea351710ae9c7d3d2426d54c950f506e

Observation 0f7c03a4-2614-492e-96a3-577d3af0b60c · inbound

Watch, Remember, Reason: Human-View Video Understanding with MLLMs cites this paper.

Watch, Remember, Reason: Human-View Video Understanding with MLLMs LLaMA-Omni2: LLM-based Real-time Spoken Chatbot with Autoregressive Streaming Speech Synthesis

Reference 127

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:27:14.747442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T22:00:28.350003Z digest=sha256:fedb3c61bcd881ba1bf7be1bd7edeca78d45c109ec4af34b7041debb8434f83c

Observation 1347e7a1-5d37-48c1-9d6b-1b492fd9f58d · inbound

Does Translation-Enhanced Speech Encoder Pre-training Affect Speech LLMs? cites this paper.

Does Translation-Enhanced Speech Encoder Pre-training Affect Speech LLMs? LLaMA-Omni2: LLM-based Real-time Spoken Chatbot with Autoregressive Streaming Speech Synthesis

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-07-04T20:30:08.189345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-25T20:03:10.858349Z digest=sha256:3bc44be24e41de77b4c66fd4c549b9958ae50ecc443859dee4cea95b09068dbb

Observation b9d06f2e-9fcc-459b-87e8-153712bfeec4 · inbound

Preserving Speech-to-Text LLM Capabilities in Speech-to-Speech Generation cites this paper.

Preserving Speech-to-Text LLM Capabilities in Speech-to-Speech Generation LLaMA-Omni2: LLM-based Real-time Spoken Chatbot with Autoregressive Streaming Speech Synthesis

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-07-01T13:05:45.816779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-01T01:01:24.536821Z digest=sha256:5455ef6c48873de62d2e03a6fc249e65335a702e5f3292aa4bdd31c18c051a61

Observation 93186e2e-187a-4d75-a08c-e04e9e56982b · inbound

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment cites this paper.

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment LLaMA-Omni2: LLM-based Real-time Spoken Chatbot with Autoregressive Streaming Speech Synthesis

Reference 139

Resolution
unresolved
no resolver link, observed 2026-08-01T07:12:17.669270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:12:17.669270Z digest=sha256:d5e77ea54262b55369f553366d7e9e9e60de2884d394de3ee034d451262e1571