Pith. sign in

Paper Citation Record · LEDGER

A Preliminary Exploration with GPT-4o Voice Mode

As of 10 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 11 inbound Pith citation observations for arXiv:2502.09940.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.09940 v1

Coverage vector

measured 31 of 31 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T20:02:51.447632Z

measured 42 of 42 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 11 of 11 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:28:53.249841Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

31 of 31 outbound references displayed

  • verified exact0
  • verified fuzzy10
  • unresolved21
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

1
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 3d7b9238-d897-4944-9000-d0126d5d7ae5 · outbound

This paper cites Qwen2-Audio Technical Report.

A Preliminary Exploration with GPT-4o Voice Mode Qwen2-Audio Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T20:02:51.293384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:02:51.293384Z digest=sha256:43722193a81d3918ccc0b733658092cf885c4b206ca77a604a2c61aa5aea7980

Observation c7c4f458-20b9-42e6-a43b-2c1309aa8a79 · outbound

This paper cites Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models.

A Preliminary Exploration with GPT-4o Voice Mode Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T20:02:51.299697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:02:51.299697Z digest=sha256:66c4855cc44209639fbd978f9e86276f764e54ba800b2127113909bcda1ccf63

Observation 6c30636a-ffaf-49f1-8a65-cd9052182336 · outbound

This paper cites Moshi: a speech-text foundation model for real-time dialogue.

A Preliminary Exploration with GPT-4o Voice Mode Moshi: a speech-text foundation model for real-time dialogue

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T20:02:51.305293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:02:51.305293Z digest=sha256:7ff40308be9a1a92184ed4d74b04baa89d74b40109d5ca764a295c5178645321

Observation 34da2de4-288e-48f0-9fb2-dc35f5362924 · outbound

This paper cites Audio Entailment: Assessing Deductive Reasoning for Audio Understanding.

A Preliminary Exploration with GPT-4o Voice Mode Audio Entailment: Assessing Deductive Reasoning for Audio Understanding

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T20:02:51.311209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:02:51.311209Z digest=sha256:df8754487f4424d42c8e604c25137171f9a0a9ec3ae81f1906f331e66836c6ae

Observation abd63aac-4a2c-415c-a107-d79b76c623d7 · outbound

This paper cites The Llama 3 Herd of Models.

A Preliminary Exploration with GPT-4o Voice Mode The Llama 3 Herd of Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T20:02:51.316910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:02:51.316910Z digest=sha256:01fd64313eec3ac73bc5166a407fd4efcd55a6e28ece5dbda507fccaea94bc45

Observation d54a85f0-f0cb-403c-84be-bf3409720cca · outbound

This paper cites LLaMA-Omni: Seamless Speech Interaction with Large Language Models.

A Preliminary Exploration with GPT-4o Voice Mode LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T20:02:51.322585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:02:51.322585Z digest=sha256:ad701b55ea3a6c7915db36f2658a6384f74e1e871760262fc504f6d796a6db81

Observation b680d517-637c-45a4-abb2-09fcb9126e8c · outbound

This paper cites WavLLM: Towards Robust and Adaptive Speech Large Language Model.

A Preliminary Exploration with GPT-4o Voice Mode WavLLM: Towards Robust and Adaptive Speech Large Language Model

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T20:02:51.327669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:02:51.327669Z digest=sha256:dbfcfa82f0c98b8880a5a026de3eabbe54ec0854374c89ba2b9bf74dd16924fd

Observation effb8c2d-a8ee-4f4b-ad31-ee6a915cf8df · outbound

This paper cites Dynamic-SUPERB Phase-2: A Collaboratively Expanding Benchmark for Measuring the Capabilities of Spoken Language Models with 180 Tasks.

A Preliminary Exploration with GPT-4o Voice Mode Dynamic-SUPERB Phase-2: A Collaboratively Expanding Benchmark for Measuring the Capabilities of Spoken Language Models with 180 Tasks

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T20:02:51.333001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:02:51.333001Z digest=sha256:a82664eaad4d386f51f811ea6d7330f50ef4bbea66b2567928ca74eff8e883a7

Observation 3bdbe4d1-afcd-4684-ba4e-9453112f679c · outbound

This paper cites Dynamic-superb: Towards a dynamic, collaborative, and comprehensive instruction-tuning benchmark for speech.

A Preliminary Exploration with GPT-4o Voice Mode Dynamic-superb: Towards a dynamic, collaborative, and comprehensive instruction-tuning benchmark for speech

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T20:02:51.337588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:02:51.337588Z digest=sha256:f0c0d32d409e0a20f31469bad8c7a7c6a816bce3d115b393a00bb368882111cf

Observation a183cfae-44ac-4211-a996-8cb3c8c1d6ec · outbound

This paper cites Audiogpt: Understanding and generating speech, music, sound, and talking head.

A Preliminary Exploration with GPT-4o Voice Mode Audiogpt: Understanding and generating speech, music, sound, and talking head

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T20:02:51.967440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T20:02:51.341839Z digest=sha256:f2e175f0e54ecabdefbb6580300ba1c0c33225eba2b18922646edc32ea3fd57c

Observation c3dc2f48-93e4-4aa5-8bfd-0802688e9f02 · outbound

This paper cites Mistral 7B.

A Preliminary Exploration with GPT-4o Voice Mode Mistral 7B

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T20:02:51.346414Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:02:51.346414Z digest=sha256:e4d26de713d71c82ee6968a9b165e2391e0b1be7caba777c77b019841454ebdb

Observation a1466721-f129-4e76-8e47-4798cfe722b4 · outbound

This paper cites Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis.

A Preliminary Exploration with GPT-4o Voice Mode Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T20:02:51.948517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T20:02:51.351641Z digest=sha256:0b08f550d8751afdae1d84bf6dfbe3143f426646712e974882a44ee108cf441f

Observation a0e53dda-3bef-4995-a635-8299557f7578 · outbound

This paper cites Speech-copilot: Leveraging large language models for speech processing via task decomposition, modularization, and program generation.

A Preliminary Exploration with GPT-4o Voice Mode Speech-copilot: Leveraging large language models for speech processing via task decomposition, modularization, and program generation

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T20:02:51.932030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T20:02:51.358132Z digest=sha256:56cb7a857b1aaa14c854555ac019de99472c1412967d548311ffdabb5235e182

Observation 34d53109-097c-4089-a55a-918649780fc9 · outbound

This paper cites Rlaif vs.

A Preliminary Exploration with GPT-4o Voice Mode Rlaif vs

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T20:02:51.915608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T20:02:51.362848Z digest=sha256:b1b21a7a8bf9881dde8af33114b3d1a2c73df200532b25f02d0ff3fe74e4702d

Observation 459e9243-efca-474f-a464-1109ebd917f7 · outbound

This paper cites The Curse of Multi-Modalities: Evaluating Hallucinations of Large Multimodal Models across Language, Visual, and Audio.

A Preliminary Exploration with GPT-4o Voice Mode The Curse of Multi-Modalities: Evaluating Hallucinations of Large Multimodal Models across Language, Visual, and Audio

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T20:02:51.368119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:02:51.368119Z digest=sha256:4eca325bb77f5c6f83a187ea3e44bf23a64a0ae8fb81e0aa90f44fb64f0c7df5

Observation 9004b082-8752-4404-b008-c3399c7284ce · outbound

This paper cites Align-SLM: Textless Spoken Language Models with Reinforcement Learning from AI Feedback.

A Preliminary Exploration with GPT-4o Voice Mode Align-SLM: Textless Spoken Language Models with Reinforcement Learning from AI Feedback

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T20:02:51.373586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:02:51.373586Z digest=sha256:6fa8e7c13f416d0246853ca92a72f6c59ecfe112b9db99ec3ed58fa74149f094

Observation 84a253a0-e7b3-4268-8b80-ebdd7e881cf8 · outbound

This paper cites Music understand- ing llama: Advancing text-to-music generation with question answering and captioning.

A Preliminary Exploration with GPT-4o Voice Mode Music understand- ing llama: Advancing text-to-music generation with question answering and captioning

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T20:02:51.900311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T20:02:51.378576Z digest=sha256:1fa046cda53697cb3fcd664aaf171868c85fc078df1867ed1f368aa393d19c99

Observation 82df58b9-b994-42e6-82d3-15658f617028 · outbound

This paper cites Generative spoken dialogue language modeling.

A Preliminary Exploration with GPT-4o Voice Mode Generative spoken dialogue language modeling

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T20:02:51.883317Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T20:02:51.383104Z digest=sha256:af095aa98f1d8181912eac181cc8bbc8f1b8df591b9199b6aef8c06c1309f83e

Observation 0982cda7-00bd-4a80-8fda-55c2e3ca948a · outbound

This paper cites GPT-4o System Card.

A Preliminary Exploration with GPT-4o Voice Mode GPT-4o System Card

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T20:02:51.387812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:02:51.387812Z digest=sha256:114033754d5300b5133501a6c55ace0ea30c3486c63941d0abbb122bb627f0c8

Observation 756a0851-4d08-4190-87d0-abef841af903 · outbound

This paper cites Training language models to follow instructions with human feedback.

A Preliminary Exploration with GPT-4o Voice Mode Training language models to follow instructions with human feedback

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T20:02:51.392746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:02:51.392746Z digest=sha256:9e807cf3098e609f0060f86a7d4132405a9a3ef1322d819677cf978e02034643

Observation 4e2dadb6-6cfd-4124-b8f1-492cf28c22bd · outbound

This paper cites MMAU: A Massive Multi-Task Audio Understanding and Reasoning Benchmark.

A Preliminary Exploration with GPT-4o Voice Mode MMAU: A Massive Multi-Task Audio Understanding and Reasoning Benchmark

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T20:02:51.397575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:02:51.397575Z digest=sha256:7cd643359fa4ffe93ebb8dcdcf51bc8729a9964581c330a28c007b0b2001f685

Observation bfe556de-7c3b-4923-b474-696c96ee8271 · outbound

This paper cites Salmonn: Towards generic hearing abilities for large language models.

A Preliminary Exploration with GPT-4o Voice Mode Salmonn: Towards generic hearing abilities for large language models

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T20:02:51.856487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T20:02:51.402731Z digest=sha256:afd675cbf5ab726399b1a084dc497733a66a6b9d1f14e0b1ea46a3b085f2e62c

Observation c8af6ec6-5620-4688-836b-3fb8d915d8ea · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

A Preliminary Exploration with GPT-4o Voice Mode Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T20:02:51.407438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:02:51.407438Z digest=sha256:3c11ea7bf2ac0ca3bfe21e5194124aba49760f55853dadb9362db1f099b388a7

Observation 66ae2e73-8b63-42ed-8ed6-697fe324dfca · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

A Preliminary Exploration with GPT-4o Voice Mode LLaMA: Open and Efficient Foundation Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T20:02:51.412615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:02:51.412615Z digest=sha256:7a1d4a479adeccc2d2c6f9d03fae1a915f8bf62d6b3dc2b24cf3a0fa2b454a67

Observation 268babd2-4d67-45e0-8765-793c569cb80c · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

A Preliminary Exploration with GPT-4o Voice Mode Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T20:02:51.417456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:02:51.417456Z digest=sha256:ec1da5d6f469686fb1b953c6623b2da71e608c82ee0064206028e7d6541df578

Observation a3bfebef-9fda-4a34-99d3-064e69857a8e · outbound

This paper cites Hear: Holistic evaluation of audio representations.

A Preliminary Exploration with GPT-4o Voice Mode Hear: Holistic evaluation of audio representations

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T20:02:51.840052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T20:02:51.422484Z digest=sha256:0147af7fcf7f50e47d60c3f5dfc5159a3c0703a97f6d7c73c92c1ca786c69332

Observation 76c2b940-f58a-4690-9095-1e6e74d0fdd6 · outbound

This paper cites Lin, Andy T.

A Preliminary Exploration with GPT-4o Voice Mode Lin, Andy T

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T20:02:51.824406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T20:02:51.427417Z digest=sha256:e51a961e3629e0685a5711dfb30c4464c8bd82c4c1c0acff8ca32397f47a3e50

Observation f7532d53-d73b-4e03-81d0-3dd53e235e36 · outbound

This paper cites Mini-Omni: Language Models Can Hear, Talk While Thinking in Streaming.

A Preliminary Exploration with GPT-4o Voice Mode Mini-Omni: Language Models Can Hear, Talk While Thinking in Streaming

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T20:02:51.432863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:02:51.432863Z digest=sha256:835d35b516712308ca78617abf792c29d04e442176e4c7295aab2d6f77e90ce2

Observation df5261ba-9ff7-49f2-87df-6dfbafda82cc · outbound

This paper cites Marble: Music audio representation benchmark for universal evaluation.

A Preliminary Exploration with GPT-4o Voice Mode Marble: Music audio representation benchmark for universal evaluation

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T20:02:51.807163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T20:02:51.437832Z digest=sha256:abf50da16148a67a2e75be863ffe3b926858cd999310db57061e1170bd053099

Observation 0d2d38f4-6f1f-45b4-a054-2146e245cb87 · outbound

This paper cites SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities.

A Preliminary Exploration with GPT-4o Voice Mode SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T20:02:51.442638Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:02:51.442638Z digest=sha256:5d3f3c70feb1a97f3374c4df4b4657e0eb7dca936ee3091acf19ad57efb5074b

Observation bf9f291e-1e32-4e63-8d09-6d3267b7679b · outbound

This paper cites SpeechAlign: Aligning Speech Generation to Human Preferences.

A Preliminary Exploration with GPT-4o Voice Mode SpeechAlign: Aligning Speech Generation to Human Preferences

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T20:02:51.447632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:02:51.447632Z digest=sha256:5d6d996becaa5ba91a3d7ee26b6dcd817c68ee104218940e7f1f6fb2346d38e9

Pith citing papers

Observation 0222dac1-65b2-4a42-b3f4-ba5a3b496eda · inbound

Breaking the Barriers of Text-Hungry and Audio-Deficient AI cites this paper.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI A Preliminary Exploration with GPT-4o Voice Mode

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:53.249841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:53.249841Z digest=sha256:9a4ac8fd8bb9a2b0759938eb16a1ca73f994ca8d696855f58dc8ada477d07b1f

Observation 63a5ae1d-f3b6-4eb2-8f1b-02d5633e1e7e · inbound

AudioLens: A Closer Look at Auditory Attribute Perception of Large Audio-Language Models cites this paper.

AudioLens: A Closer Look at Auditory Attribute Perception of Large Audio-Language Models A Preliminary Exploration with GPT-4o Voice Mode

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:40.942515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:40.942515Z digest=sha256:1db2af4196cfef37471d04295837ce883f352c5aed1e596e56202140b6ae21e1

Observation bab67a9a-cbc5-4ae2-a27f-821dd93357a5 · inbound

Towards Generalized Source Tracing for Codec-Based Deepfake Speech cites this paper.

Towards Generalized Source Tracing for Codec-Based Deepfake Speech A Preliminary Exploration with GPT-4o Voice Mode

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:08.435473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:08.435473Z digest=sha256:fd4e780321662d2da4ecadd27fbe394ea63378bdb30bf9fb728c6cf1ae46b7e3

Observation 7e20f3cb-0194-4c2e-9f4e-30aa19a240bd · inbound

A Survey of Automatic Evaluation Methods on Text, Visual and Speech Generations cites this paper.

A Survey of Automatic Evaluation Methods on Text, Visual and Speech Generations A Preliminary Exploration with GPT-4o Voice Mode

Reference 208

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:46.980945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:46.980945Z digest=sha256:714b92d28d3f191b24892140f010dcf9955e76fe5c5de75535ee517d7e24d5a3

Observation 68ec82f6-b36f-4b8a-a48c-6ad95da979b8 · inbound

When Silence Matters: The Impact of Irrelevant Audio on Text Reasoning in Large Audio-Language Models cites this paper.

When Silence Matters: The Impact of Irrelevant Audio on Text Reasoning in Large Audio-Language Models A Preliminary Exploration with GPT-4o Voice Mode

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-18T11:11:17.870042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-18T11:08:08.916893Z digest=sha256:8345aa22a14feae12d1b52aaea3b0f31f03d335ab01ad2d57bb781ab8e2a26f4

Observation fa56d9ba-6427-47f9-997d-c19382bb01b9 · inbound

Style Amnesia: Investigating Speaking Style Degradation and Mitigation in Multi-Turn Spoken Language Models cites this paper.

Style Amnesia: Investigating Speaking Style Degradation and Mitigation in Multi-Turn Spoken Language Models A Preliminary Exploration with GPT-4o Voice Mode

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-16T19:28:20.154420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-16T19:26:30.384474Z digest=sha256:592667473d9a396328a7ffdd1ddd3c07dd41f4a263c0eebb87cf357ec838f15e

Observation b4d92f47-b149-4acb-bda9-9f8c2e7888d7 · inbound

ASPIRin: Action Space Projection for Interactivity-Optimized Reinforcement Learning in Full-Duplex Speech Language Models cites this paper.

ASPIRin: Action Space Projection for Interactivity-Optimized Reinforcement Learning in Full-Duplex Speech Language Models A Preliminary Exploration with GPT-4o Voice Mode

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-11T07:35:58.745428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T17:05:45.214298Z digest=sha256:77eb026163539a8c2db8f51d288f2b4b8c91119b25001533f40119e5c12157c8

Observation 2a7258c4-95a8-4b3f-9e9f-a4f17429687a · inbound

All That Glitters Is Not Audio: Rethinking Text Priors and Audio Reliance in Audio-Language Evaluation cites this paper.

All That Glitters Is Not Audio: Rethinking Text Priors and Audio Reliance in Audio-Language Evaluation A Preliminary Exploration with GPT-4o Voice Mode

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:16:36.175798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-07T17:39:38.234052Z digest=sha256:f4a6fa48210f051ea44ba9123ed107471606b9c5f7d5bd2ee4556252d4ddd625

Observation 8f148614-e12b-440d-8fc7-495f8f1f40f2 · inbound

ISCSLP 2026 CoT-TTS Challenge: Chain-of-Thought Reasoning for Context-Aware Text-to-Speech cites this paper.

ISCSLP 2026 CoT-TTS Challenge: Chain-of-Thought Reasoning for Context-Aware Text-to-Speech A Preliminary Exploration with GPT-4o Voice Mode

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-07-04T08:29:42.197627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-26T11:33:32.067771Z digest=sha256:5e75a1ad59cda3bb0996e9b4529814bdd74e45d5e98ad0f6e589fbb0962abcad

Observation 21f6f1e5-7025-4737-8c84-c0a4aed49abd · inbound

Hierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMs cites this paper.

Hierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMs A Preliminary Exploration with GPT-4o Voice Mode

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-07-08T02:44:27.576608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-07-08T02:38:31.073805Z digest=sha256:888d128cf7ac80bc0647f352a8947b89881ad5ff4442d57fa3cdfbd9c083d4ad

Observation b40423e5-fe69-42c1-84b8-3cdb4d8f8380 · inbound

Encoder-Side Neuron Identification and Amplification for Acoustic Perception in Large Audio-Language Models cites this paper.

Encoder-Side Neuron Identification and Amplification for Acoustic Perception in Large Audio-Language Models A Preliminary Exploration with GPT-4o Voice Mode

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-14T03:05:43.100032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T03:05:43.100032Z digest=sha256:705a702d3505f4178d09df27bafd58e276bfac6d61b544019ce01373314d183a