Pith. sign in

Paper Citation Record · LEDGER

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model

As of 8 August 2026, this Paper Citation Record lists 55 of 55 outbound references and 4 inbound Pith citation observations for arXiv:2506.08967.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.08967 v2

Coverage vector

measured 55 of 55 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:03:16.518476Z

measured 59 of 59 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-16T05:59:50.900436Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-16T05:59:51.133845Z

Reference resolution

55 of 55 outbound references displayed

  • verified exact0
  • verified fuzzy18
  • unresolved37
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 10cb4c00-fb42-4a17-8643-93c8cb711898 · outbound

This paper cites Claude 3.5 sonnet.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model Claude 3.5 sonnet

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:03:20.303920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:03:11.339879Z digest=sha256:898de2ccc25f9447916a108f172ddec48463ade294a33a8ec6403ab949b479ed

Observation abd46843-edbf-447c-9d90-54b740e55788 · outbound

This paper cites ML4CO-KIDA: Knowledge Inheritance in Dataset Aggregation.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model ML4CO-KIDA: Knowledge Inheritance in Dataset Aggregation

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:11.413339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:11.413339Z digest=sha256:06fb90a7f22c162258046d1dc7f20b97e2e47cd9feb44123779f22b13c9a1272

Observation c04f6727-a3cd-412a-8844-60e880dc20d2 · outbound

This paper cites VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:11.529192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:11.529192Z digest=sha256:5fbdbaa88b970d6a84dab70bd881f6a03956d3c8513f3e73c0aa96a4d7b6de89

Observation d0a5322d-cb5c-40cb-a5c9-1180a7413fd3 · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:11.622724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:11.622724Z digest=sha256:706b4ff235caf0bfe65359af11d0bb9bc22978a99a62a909ac94c3b267024f7f

Observation a22e50bd-1de3-49eb-aca5-215897c7b636 · outbound

This paper cites Qwen2-Audio Technical Report.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model Qwen2-Audio Technical Report

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:11.715744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:11.715744Z digest=sha256:b261697aac6749cd08a6cd3bb720e40036167ce6545a7832039793671dfa456e

Observation 6c0ef7c8-88f3-49f6-a9a9-767451b04f45 · outbound

This paper cites Simple and controllable music generation.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model Simple and controllable music generation

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:03:20.146037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:03:11.799823Z digest=sha256:d875cc1ed505bb19f0ec7bc114b54d8ca75966ab417987eae77be23d189e5f94

Observation 41695169-4cb5-48d0-8092-c55a769324ee · outbound

This paper cites Recent Advances in Speech Language Models: A Survey.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model Recent Advances in Speech Language Models: A Survey

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:11.869766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:11.869766Z digest=sha256:96aa2c43aab08dec3c283cf120e490c4a2ee9dbc5e63e605cc777564759075c1

Observation 98a60a68-68ec-4beb-b36f-3992a44c7334 · outbound

This paper cites Pengi: An audio language model for audio tasks.Advances in Neural Information Processing Systems, 36:18090–18108, 2023.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model Pengi: An audio language model for audio tasks.Advances in Neural Information Processing Systems, 36:18090–18108, 2023

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:03:19.991060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:03:11.955150Z digest=sha256:616fd989e2e2626ef22c1f37ad3192c63c5583fb335efffac00c1db64e3c6e03

Observation 04ae4ddb-c8ae-4042-b533-bbee52d4e131 · outbound

This paper cites Kimi-Audio Technical Report.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model Kimi-Audio Technical Report

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:12.052740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:12.052740Z digest=sha256:1ed27f6699625b4f873ab32522f6ced45f4998c147a14c51b3af4f5e6940da88

Observation 054ed7e2-d761-4403-9623-b840e7a9ba1e · outbound

This paper cites CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:12.174526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:12.174526Z digest=sha256:7beb6fd2cb2591a9bfabe1c6f76a589c139eb90675bf926cd0eabb1ce9707020

Observation 1daf89f6-b755-43eb-ba41-ef48f608a2c6 · outbound

This paper cites Emo-dpo: Controllable emo- tional speech synthesis through direct preference optimization.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model Emo-dpo: Controllable emo- tional speech synthesis through direct preference optimization

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:03:19.785059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:03:12.244985Z digest=sha256:6dc75b6d3d2a6c6f5acef1706729ef32f356a07e8f54f680af377823b9f4f929

Observation debb1561-7cd4-450d-94b9-44605811f8d0 · outbound

This paper cites Paraformer: Fast and Accurate Parallel Transformer for Non-autoregressive End-to-End Speech Recognition.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model Paraformer: Fast and Accurate Parallel Transformer for Non-autoregressive End-to-End Speech Recognition

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:12.369009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:12.369009Z digest=sha256:d90cb0b5eb7e324b33b4ad220f7211e88841089a5eae6714ccd99cc021001872

Observation a942c027-9629-45eb-a593-242adb83c897 · outbound

This paper cites Joint audio and speech understanding.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model Joint audio and speech understanding

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:03:19.651919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:03:12.467905Z digest=sha256:32e309b44fdf8ebfb951d5f9a79207921d6483be990dffd4ce490ea2b51da81d

Observation b205fbe5-d36f-4836-bd3b-f5605229ace5 · outbound

This paper cites Gemini 2.0 pro.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model Gemini 2.0 pro

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:03:19.537131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:03:12.579313Z digest=sha256:b0baff3cd5b4d339e145d4dff7ba78c03fc6e4fc4cf787f7509d89a499fc130c

Observation f127b7f0-c905-4dec-8d96-fee8ce8bed7d · outbound

This paper cites The Llama 3 Herd of Models.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model The Llama 3 Herd of Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:12.674168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:12.674168Z digest=sha256:5dff6f6501b85b727194b27f1d2b9db884451f3c883509700d55d50273b63eef

Observation 6d71bed5-e07f-464b-8b25-890a7a7847c6 · outbound

This paper cites VALL-E R: Robust and Efficient Zero-Shot Text-to-Speech Synthesis via Monotonic Alignment.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model VALL-E R: Robust and Efficient Zero-Shot Text-to-Speech Synthesis via Monotonic Alignment

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:12.752879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:12.752879Z digest=sha256:c4b0cc338c63e03eaade4fa32c39b083c5f5512ccdea2d3d67e4ffcb210922b2

Observation fa979691-538a-4d86-a543-e88cc561bffa · outbound

This paper cites Deep residual learning for image recognition.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model Deep residual learning for image recognition

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:03:19.402351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:03:12.827575Z digest=sha256:3f05437270cb887023aa59c6b2f64865250b18ef2c4503b96e7dab3601beec86

Observation 82e4a2a7-0850-40ef-9200-49d35d6dd197 · outbound

This paper cites DrivingWorld: Constructing World Model for Autonomous Driving via Video GPT.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model DrivingWorld: Constructing World Model for Autonomous Driving via Video GPT

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:12.935432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:12.935432Z digest=sha256:774c84e7fad02eaef58c30723bc87d9af8617c03412bfae0065cd3f6165c8c98

Observation 2ecf714d-a632-43b9-948e-0244e0216d2f · outbound

This paper cites Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:13.012840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:13.012840Z digest=sha256:2a2d4c2fc4764f6b591645067f886c6c21ed57eeba3a972a6577510180db0614

Observation e0ab10ad-e0c6-4cc2-9cb0-80ac60a65f79 · outbound

This paper cites Audiogpt: Understanding and generating speech, music, sound, and talking head.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model Audiogpt: Understanding and generating speech, music, sound, and talking head

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:13.099690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:13.099690Z digest=sha256:1c9aadbee88cda423749af8e428862d2afbc81b31f7b1a13a881482358e6c1ee

Observation 4f417131-a65b-427e-8cfc-91cb7a895586 · outbound

This paper cites GPT-4o System Card.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model GPT-4o System Card

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:13.194856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:13.194856Z digest=sha256:e35b7eb63ad4e300f6d7ed66476ae69231f04d1a698af36b0eca0e78fb714c10

Observation 8dfec203-8715-435f-9432-4cfbb285ceb1 · outbound

This paper cites OpenAI o1 System Card.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model OpenAI o1 System Card

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:13.295965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:13.295965Z digest=sha256:f4521c882702d669b0a8ce82610402f7152095708e4e50b10132fd2a52c3d5d7

Observation 681a5cdb-7ba7-4382-9e1f-b936d7d04130 · outbound

This paper cites WavChat: A Survey of Spoken Dialogue Models.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model WavChat: A Survey of Spoken Dialogue Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:13.393147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:13.393147Z digest=sha256:46a3549dbf5db4c06d54967b0a0c2f17c3583e1d5d71577d67cec67cf9a76d5d

Observation 3b51b3cd-7c60-4af5-b140-ae3258ec7077 · outbound

This paper cites An llm compiler for parallel function calling.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model An llm compiler for parallel function calling

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:03:19.310225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:03:13.488079Z digest=sha256:d4e44cd76101838cd8cc64c6489b439b16b73e988671baab8e29256c10e2c1ee

Observation 4784308c-b4a6-411a-855a-1dc0db1d4d6d · outbound

This paper cites Understanding the Effects of RLHF on LLM Generalisation and Diversity.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:13.573676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:13.573676Z digest=sha256:e9f4dfa55791ed4236fe80586a98dbe461b8cb9c54d2ed434a8b5f1b5b181127

Observation 4d3e3944-1c43-49b2-a452-e36b1b2eda44 · outbound

This paper cites Audio Flamingo: A Novel Audio Language Model with Few-Shot Learning and Dialogue Abilities.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model Audio Flamingo: A Novel Audio Language Model with Few-Shot Learning and Dialogue Abilities

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:13.651263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:13.651263Z digest=sha256:0134391a2f3184be6cd3bff3f83a39485cbcd73f142e4d82d4e99153468c644b

Observation fa47b73e-e297-4d3f-a577-85bf14e82af4 · outbound

This paper cites Flexkbqa: A flexible llm-powered framework for few-shot knowledge base question answering.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model Flexkbqa: A flexible llm-powered framework for few-shot knowledge base question answering

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:03:19.137327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:03:13.754388Z digest=sha256:707bb670c0a89c5e5d8ae4363c4042a185b57ac1992801b9dc1ed851bf5af1b6

Observation 945b7c88-6470-46fa-989f-fd6bb4e79769 · outbound

This paper cites Merging models with fisher-weighted averaging.Advances in Neural Information Processing Systems, 35:17703–17716, 2022.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model Merging models with fisher-weighted averaging.Advances in Neural Information Processing Systems, 35:17703–17716, 2022

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:13.853463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:13.853463Z digest=sha256:9b5ed42477701e1dfcb2ab5412c72afde1a8bc3b0d51166d6a25cbe7232f7543

Observation a3ac23ad-a909-40b9-81fe-b2e8d0973b32 · outbound

This paper cites Using an llm to help with code understanding.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model Using an llm to help with code understanding

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:13.950462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:13.950462Z digest=sha256:4fa244ab3beae8a3f453cc5e1e15e69e05563f9871f494a36ad071c31741e674

Observation 6c60b368-3315-473a-bc21-6a6e4f930483 · outbound

This paper cites The role of paralinguistic cues in social life.ANALYSIS OF MODERN SCIENCE AND INNOVATION, 1(2):215–218, 2024.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model The role of paralinguistic cues in social life.ANALYSIS OF MODERN SCIENCE AND INNOVATION, 1(2):215–218, 2024

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:03:18.931107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:03:14.045361Z digest=sha256:15d84b6283ce76eb092363f19506893f5097440aee71af43ffc4240b575be239

Observation 2905d971-6f5a-4ca6-b4ba-96df5210b1ce · outbound

This paper cites A survey on speech large language models.arXiv preprint arXiv:2410.18908, 2024.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model A survey on speech large language models.arXiv preprint arXiv:2410.18908, 2024

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:14.141710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:14.141710Z digest=sha256:ffe163db12057b6e0e1e3f37beb6a83f9f6701ba6f00a17f0541cd2f4f88bc9c

Observation cd1534b0-420e-4502-8a87-676fc3828313 · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.Advances in Neural Information Processing Systems, 36:53728–53741, 2023.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model Direct preference optimization: Your language model is secretly a reward model.Advances in Neural Information Processing Systems, 36:53728–53741, 2023

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:14.234520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:14.234520Z digest=sha256:2f072f467a2420eebfa3b055ffd4538c963f57680039127c9b143ec40b91596d

Observation 7f2b0cb2-a3ab-483b-adf7-e8e47cc6535a · outbound

This paper cites AudioPaLM: A Large Language Model That Can Speak and Listen.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model AudioPaLM: A Large Language Model That Can Speak and Listen

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:14.333366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:14.333366Z digest=sha256:b5b7893cbd32230f6fa2ab7b1ec348ce6cdf6e096be0b421275cc0b6007f9938

Observation ddd32cd2-70b9-49a6-85db-d6918228071c · outbound

This paper cites How to debug code with github copilot.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model How to debug code with github copilot

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:03:18.762496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:03:14.429648Z digest=sha256:876e0ef2564547866bf928aa2bc22c186a9afdf6ea37c4518b4ffb7f191fe2b4

Observation 0450f964-3514-4b32-b6d4-1193c5f26a67 · outbound

This paper cites Paralinguistics in speech and language—state-of-the-art and the challenge.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model Paralinguistics in speech and language—state-of-the-art and the challenge

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:03:18.596275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:03:14.503974Z digest=sha256:e3bb673176940aafc3d18307559a99330ddaf5afc61778be9807a9cb340bfcdd

Observation 1f9684bb-3fd8-44ed-ba66-efc89785ae4c · outbound

This paper cites HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in Hugging Face.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in Hugging Face

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:14.618551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:14.618551Z digest=sha256:e1c1bb8197526945dfee41807e970698b5bcd54889dcca0d51cc118b00df7b83

Observation 2c510c4f-08a8-4eb0-a358-fcd62f230dd2 · outbound

This paper cites Stepeval-audio-360.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model Stepeval-audio-360

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:03:18.443217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:03:14.752046Z digest=sha256:2b35b3cafa5c88e20c970e93fb78757bf1e1b7bb45176e034146783a04e2fa58

Observation 95470915-60f4-4c1e-8152-ab198d7116a5 · outbound

This paper cites SALMONN: Towards Generic Hearing Abilities for Large Language Models.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model SALMONN: Towards Generic Hearing Abilities for Large Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:14.876677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:14.876677Z digest=sha256:9dd81653f2a38419479f882377ff8bca8ecbef590cfe68aed06637a6f6f771c1

Observation 81d875f8-529e-42f4-811b-7d2830f9f202 · outbound

This paper cites Preference alignment improves language model-based tts.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model Preference alignment improves language model-based tts

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:03:18.297584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:03:14.945288Z digest=sha256:11ec1bd72e6e511e884db680eadac7198846d54bd8be249b4e8489ec0251fa9a

Observation d1195625-63a9-4c2a-b0c9-f18331437a73 · outbound

This paper cites Lami: Large language models for multi-modal human-robot interaction.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model Lami: Large language models for multi-modal human-robot interaction

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:03:18.113523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:03:15.071882Z digest=sha256:e3998a57c7dd87f85b46f02b26a2a0f2108014e931067599e674373dad52c14c

Observation 907f9a29-8e0a-4a27-99f6-cd12850217ed · outbound

This paper cites Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:15.171753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:15.171753Z digest=sha256:ce77882db2c723dc8d187a3e88c8cd76b99dbecf04f591bf31c1bf24c24f0664

Observation 41321313-bb77-41b8-856f-84ad7fcbe144 · outbound

This paper cites Reinforcement Learning for LLM Post-Training: A Survey.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model Reinforcement Learning for LLM Post-Training: A Survey

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:15.250044Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:15.250044Z digest=sha256:3948537a340013cd50bc793861251129d9a2bfce47e85220047073e0a3c8fce4

Observation 66595b68-8294-47cb-94a7-89faf54611fd · outbound

This paper cites Model soups: aver- aging weights of multiple fine-tuned models improves accuracy without increasing inference time.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model Model soups: aver- aging weights of multiple fine-tuned models improves accuracy without increasing inference time

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:15.331162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:15.331162Z digest=sha256:9a53c92f1a387a3189b7311e2b516389b5873445a4de090f2ffe92c05c9a9af8

Observation 962da2ef-045a-4b6e-b459-fe49ff99f826 · outbound

This paper cites Codec-SUPERB: An In-Depth Analysis of Sound Codec Models.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model Codec-SUPERB: An In-Depth Analysis of Sound Codec Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:15.425974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:15.425974Z digest=sha256:f1091d0162c7ff83fe686d87552764d6a8dcad734be3dfb199315b2406c5154f

Observation 07125a7f-1975-4d32-a803-21d414475f7b · outbound

This paper cites When search engine services meet large language models: visions and challenges.IEEE Transactions on Services Computing, 2024.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model When search engine services meet large language models: visions and challenges.IEEE Transactions on Services Computing, 2024

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:03:17.895203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:03:15.557323Z digest=sha256:23d4947fa70f574aeaded9a7ad1a20a7df929c6c5ece1139d3f5c3a2e7c51683

Observation 7f1bb81c-d213-49f6-b752-269636b914d0 · outbound

This paper cites Qwen2.5-Omni Technical Report.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model Qwen2.5-Omni Technical Report

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:15.651302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:15.651302Z digest=sha256:2c7967b752805f81ac5ef4ffc905f2f542fc302ce51efeb532e24bb4dbc939a6

Observation bd288521-e2f9-45f7-b906-647636d89695 · outbound

This paper cites Uniaudio: Towards universal audio generation with large language models.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model Uniaudio: Towards universal audio generation with large language models

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:03:17.676913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:03:15.749688Z digest=sha256:c6876294864dd0f0c84872eec184bf2c69e40a18baf1918fb7b802dfbbddfe2e

Observation 5eade042-b8c6-4f85-aa5f-faaa4851b64b · outbound

This paper cites Soundstream: An end-to-end neural audio codec.IEEE/ACM Transactions on Audio, Speech, and Language Processing, 30:495–507, 2021.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model Soundstream: An end-to-end neural audio codec.IEEE/ACM Transactions on Audio, Speech, and Language Processing, 30:495–507, 2021

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:15.837924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:15.837924Z digest=sha256:26b487357835adc7333eb67ceed1c154b59064d6c72d51d9d5da6c8a5be69db1

Observation 6364a05c-56a0-4e4d-a795-8ba0b4bfed1d · outbound

This paper cites GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:15.913730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:15.913730Z digest=sha256:5675e38e72c5995ed0c2803183985aaa7a313eeeadeb8d78b26ea43991331c98

Observation fb0bdad5-184f-4cb4-a36d-75ce3e54ac98 · outbound

This paper cites Root mean square layer normalization.Advances in Neural Information Processing Systems, 32, 2019.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model Root mean square layer normalization.Advances in Neural Information Processing Systems, 32, 2019

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:15.976166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:15.976166Z digest=sha256:7399af929add1265874215950121b07b21779eda94f6e63e080fb0745ba5fbbe

Observation 0d49440e-e10f-4e6a-9510-fbd1db9a59de · outbound

This paper cites SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:16.062120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:16.062120Z digest=sha256:c1d2d647ec86a86c8d2be54de2235a8326320cbfa591ab5397dd70f2896aa7ff

Observation 359d011e-1276-4af3-b3e7-480e3cd1ae5c · outbound

This paper cites SpeechAlign: Aligning Speech Generation to Human Preferences.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model SpeechAlign: Aligning Speech Generation to Human Preferences

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:16.158400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:16.158400Z digest=sha256:b5ea4256071cd739ebaf5c0b7fdcb7eb160551cff875b9164ae5c0199f257736

Observation 2e36a9ec-f727-4c09-a3e8-6fcb7d392395 · outbound

This paper cites Vistorybench: Comprehensive benchmark suite for story visualiza- tion.arXiv preprint arXiv:2505.24862, 2025.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model Vistorybench: Comprehensive benchmark suite for story visualiza- tion.arXiv preprint arXiv:2505.24862, 2025

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:16.265675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:16.265675Z digest=sha256:17787a36579d17d1845ceda95ccedce784a6841e9a02c8413ad57225ce30a5c2

Observation 4e967e3e-1770-489c-9c03-e9a1b3228968 · outbound

This paper cites Toolqa: A dataset for llm question answering with external tools.Advances in Neural Information Processing Systems, 36:50117–50143, 2023.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model Toolqa: A dataset for llm question answering with external tools.Advances in Neural Information Processing Systems, 36:50117–50143, 2023

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:16.429765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:16.429765Z digest=sha256:150fd1a40db89c3b333ff3d90ea6b5f7c133e30353e9dc14d074335e5c0a3715

Observation a601aaa1-bf4b-4671-b545-3f298fe43b31 · outbound

This paper cites Pre-trained language model based ranking in baidu search.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model Pre-trained language model based ranking in baidu search

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:03:17.385454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:03:16.518476Z digest=sha256:ec1da5c59d77816601a65737a87ebc9ace07ae688b5bf7abdaf6acf8a26d1255

Pith citing papers

Observation 77ab2a72-895c-410e-9eb3-653860b0be7d · inbound

Step-Audio 2 Technical Report cites this paper.

Step-Audio 2 Technical Report Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-16T05:59:51.136199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T05:59:50.900436Z digest=sha256:7feb6db840dfb98a9ea0e56225ab02b491ba401208832f61e1028e89b001348c

Observation ac1ddd35-444d-4687-88c5-1ac7ceb4cc4e · inbound

WavAlign: Enhancing Intelligence and Expressiveness in Spoken Dialogue Models via Adaptive Hybrid Post-Training cites this paper.

WavAlign: Enhancing Intelligence and Expressiveness in Spoken Dialogue Models via Adaptive Hybrid Post-Training Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T11:10:09.183579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T11:06:07.355633Z digest=sha256:3fa344797b41981f8a6ad820617b1360f4ede805876a8ab0f1eaff25893a6939

Observation 86c944ae-9249-46eb-8c34-1251a9201072 · inbound

Aligning Language Models for Lyric-to-Melody Generation with Rule-Based Musical Constraints cites this paper.

Aligning Language Models for Lyric-to-Melody Generation with Rule-Based Musical Constraints Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-11T12:41:05.205344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T03:09:41.001660Z digest=sha256:519e01bf32bda0fbe5a45fe6740482072f17f371f8d6e356ef39018776351c0c

Observation ab9d91c2-f249-4a97-9649-5b3a42591fff · inbound

VITA-QinYu: Expressive Spoken Language Model for Role-Playing and Singing cites this paper.

VITA-QinYu: Expressive Spoken Language Model for Role-Playing and Singing Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T04:50:56.014102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-11T01:03:09.942984Z digest=sha256:6a1d530eed660de1d37df32bf20c6067190b25a9baeec5233619d3498400db74