Pith. sign in

Paper Citation Record · LEDGER

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model

As of 18 August 2026, this Paper Citation Record lists 55 of 55 outbound references and 4 inbound Pith citation observations for arXiv:2506.08967.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.08967 v2

Coverage vector

measured 55 of 55 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:03:16.518476Z

measured 59 of 59 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-16T05:59:50.900436Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-16T05:59:51.133845Z

Reference resolution

55 of 55 outbound references displayed

  • verified exact0
  • verified fuzzy18
  • unresolved37
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 10cb4c00-fb42-4a17-8643-93c8cb711898 · outbound

This paper cites Claude 3.5 sonnet.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model Claude 3.5 sonnet

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:03:20.303920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:03:11.339879Z digest=sha256:93553553893e3a2930e7aa743911be2db19b009a6fe3c3f14b1b1ae1cb33bc48

Observation abd46843-edbf-447c-9d90-54b740e55788 · outbound

This paper cites ML4CO-KIDA: Knowledge Inheritance in Dataset Aggregation.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model ML4CO-KIDA: Knowledge Inheritance in Dataset Aggregation

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:11.413339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:11.413339Z digest=sha256:d27cb159a5bc666e650aa3032992a7902a537471e02a74cd875b8a770a0b9a86

Observation c04f6727-a3cd-412a-8844-60e880dc20d2 · outbound

This paper cites VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:11.529192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:11.529192Z digest=sha256:31f4e7b59105f06396c87e406cded500a660710e64197366d8ec03205e4a5c42

Observation d0a5322d-cb5c-40cb-a5c9-1180a7413fd3 · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:11.622724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:11.622724Z digest=sha256:c0a492eebafee752f9988bd612072ec2da61c796140add826f253a63b46da101

Observation a22e50bd-1de3-49eb-aca5-215897c7b636 · outbound

This paper cites Qwen2-Audio Technical Report.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model Qwen2-Audio Technical Report

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:11.715744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:11.715744Z digest=sha256:85a161022a452447bd8d205302f74db0f9b5d679b09eea260e9a02947ac764d8

Observation 6c0ef7c8-88f3-49f6-a9a9-767451b04f45 · outbound

This paper cites Simple and controllable music generation.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model Simple and controllable music generation

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:03:20.146037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:03:11.799823Z digest=sha256:ba46129322776ee622e2dab1aa67424d44c534264fe0ed734a7bde4ca2d8c323

Observation 41695169-4cb5-48d0-8092-c55a769324ee · outbound

This paper cites Recent Advances in Speech Language Models: A Survey.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model Recent Advances in Speech Language Models: A Survey

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:11.869766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:11.869766Z digest=sha256:b5de0be9cf2950a1bb44a7f0222f943a456b0bf7e9bf70737a4359eeb53bcae6

Observation 98a60a68-68ec-4beb-b36f-3992a44c7334 · outbound

This paper cites Pengi: An audio language model for audio tasks.Advances in Neural Information Processing Systems, 36:18090–18108, 2023.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model Pengi: An audio language model for audio tasks.Advances in Neural Information Processing Systems, 36:18090–18108, 2023

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:03:19.991060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:03:11.955150Z digest=sha256:64723d8ae92bfbe09185d619d7e6e84179f199caa9c42ff1661adddcd8d6f37b

Observation 04ae4ddb-c8ae-4042-b533-bbee52d4e131 · outbound

This paper cites Kimi-Audio Technical Report.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model Kimi-Audio Technical Report

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:12.052740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:12.052740Z digest=sha256:cd0dca0eab4ca4872db079748783abeea9aacd85f7a3921adc689418ee1f9c7f

Observation 054ed7e2-d761-4403-9623-b840e7a9ba1e · outbound

This paper cites CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:12.174526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:12.174526Z digest=sha256:92bee6d8675ccf1d4a31f39822c4905c647ece09ce2f3afd9743d2f720605244

Observation 1daf89f6-b755-43eb-ba41-ef48f608a2c6 · outbound

This paper cites Emo-dpo: Controllable emo- tional speech synthesis through direct preference optimization.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model Emo-dpo: Controllable emo- tional speech synthesis through direct preference optimization

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:03:19.785059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:03:12.244985Z digest=sha256:9f6ca1c67f01f29744d9a8f6f5c3e9cf84abc0f661dec92a3793c764b48b9fc2

Observation debb1561-7cd4-450d-94b9-44605811f8d0 · outbound

This paper cites Paraformer: Fast and Accurate Parallel Transformer for Non-autoregressive End-to-End Speech Recognition.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model Paraformer: Fast and Accurate Parallel Transformer for Non-autoregressive End-to-End Speech Recognition

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:12.369009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:12.369009Z digest=sha256:67e4eeeb3b50a7b4cd2c429f71b0fdf04a27416a99330d75ae03250df1bc1359

Observation a942c027-9629-45eb-a593-242adb83c897 · outbound

This paper cites Joint audio and speech understanding.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model Joint audio and speech understanding

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:03:19.651919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:03:12.467905Z digest=sha256:8da7809183f6873953a0b5d36dacb9907e4cb58173968d7d3817430b831f7626

Observation b205fbe5-d36f-4836-bd3b-f5605229ace5 · outbound

This paper cites Gemini 2.0 pro.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model Gemini 2.0 pro

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:03:19.537131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:03:12.579313Z digest=sha256:ef2318a53b36c815f8b92bcbe85a7bf2951eae11b9cb48e72778208306edad3e

Observation f127b7f0-c905-4dec-8d96-fee8ce8bed7d · outbound

This paper cites The Llama 3 Herd of Models.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model The Llama 3 Herd of Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:12.674168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:12.674168Z digest=sha256:8d92f7efa58fd9ab82d8fea078f41dab4c7bc1b7183ec42e94c1eeb686124753

Observation 6d71bed5-e07f-464b-8b25-890a7a7847c6 · outbound

This paper cites VALL-E R: Robust and Efficient Zero-Shot Text-to-Speech Synthesis via Monotonic Alignment.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model VALL-E R: Robust and Efficient Zero-Shot Text-to-Speech Synthesis via Monotonic Alignment

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:12.752879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:12.752879Z digest=sha256:0c33a7395c1634b0a5bc5625a72225781617618d9bcc2b139a6bf6a6fcc5719b

Observation fa979691-538a-4d86-a543-e88cc561bffa · outbound

This paper cites Deep residual learning for image recognition.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model Deep residual learning for image recognition

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:03:19.402351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:03:12.827575Z digest=sha256:0864f85ee4044c6cc49813d3ad65529e0f332e9ad026a158e8ad6343786b9a16

Observation 82e4a2a7-0850-40ef-9200-49d35d6dd197 · outbound

This paper cites DrivingWorld: Constructing World Model for Autonomous Driving via Video GPT.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model DrivingWorld: Constructing World Model for Autonomous Driving via Video GPT

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:12.935432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:12.935432Z digest=sha256:cc5fcfeacb96bdd981e93d9a2e35b21391d00a103dc7df52bdadfc9b207b70dd

Observation 2ecf714d-a632-43b9-948e-0244e0216d2f · outbound

This paper cites Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:13.012840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:13.012840Z digest=sha256:732de9889aef968c7cdf1e54144c9202b1e70c26d3e80bcf840d8d5a8ab60e40

Observation e0ab10ad-e0c6-4cc2-9cb0-80ac60a65f79 · outbound

This paper cites Audiogpt: Understanding and generating speech, music, sound, and talking head.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model Audiogpt: Understanding and generating speech, music, sound, and talking head

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:13.099690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:13.099690Z digest=sha256:69ea9520f3ec00cb4262278e4332c9d796a47beb0cbfba8e38aeb859d1851941

Observation 4f417131-a65b-427e-8cfc-91cb7a895586 · outbound

This paper cites GPT-4o System Card.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model GPT-4o System Card

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:13.194856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:13.194856Z digest=sha256:1699c30b317af40ebbf58ab173d2fd52c8e13b4108fafb07c9a1c913c79283c8

Observation 8dfec203-8715-435f-9432-4cfbb285ceb1 · outbound

This paper cites OpenAI o1 System Card.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model OpenAI o1 System Card

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:13.295965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:13.295965Z digest=sha256:f7a8d8c1d5aa73c8b81ba3f8f26481d52eda7531ea7c788a542b40b42e91edb3

Observation 681a5cdb-7ba7-4382-9e1f-b936d7d04130 · outbound

This paper cites WavChat: A Survey of Spoken Dialogue Models.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model WavChat: A Survey of Spoken Dialogue Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:13.393147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:13.393147Z digest=sha256:bc7ca2ad6ebc7331c559d4ab49b79add731d72530167ceddd42837af371f14ea

Observation 3b51b3cd-7c60-4af5-b140-ae3258ec7077 · outbound

This paper cites An llm compiler for parallel function calling.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model An llm compiler for parallel function calling

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:03:19.310225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:03:13.488079Z digest=sha256:e31d26e2fbd747eb214572a4b1de331e7b3c6b5122827c9ca2b848780a15f044

Observation 4784308c-b4a6-411a-855a-1dc0db1d4d6d · outbound

This paper cites Understanding the Effects of RLHF on LLM Generalisation and Diversity.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:13.573676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:13.573676Z digest=sha256:8a692f14ab3c9767428d311408201fb667633b54c271718c2d42d980e7609421

Observation 4d3e3944-1c43-49b2-a452-e36b1b2eda44 · outbound

This paper cites Audio Flamingo: A Novel Audio Language Model with Few-Shot Learning and Dialogue Abilities.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model Audio Flamingo: A Novel Audio Language Model with Few-Shot Learning and Dialogue Abilities

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:13.651263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:13.651263Z digest=sha256:4d5f67205f882b5487cc5431c5ac501ce99e0dd046df07f450d3d8e242b232a7

Observation fa47b73e-e297-4d3f-a577-85bf14e82af4 · outbound

This paper cites Flexkbqa: A flexible llm-powered framework for few-shot knowledge base question answering.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model Flexkbqa: A flexible llm-powered framework for few-shot knowledge base question answering

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:03:19.137327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:03:13.754388Z digest=sha256:cdcc7db9c780382ef5167071787d5652c3754371ee51d47a7567c836a8313367

Observation 945b7c88-6470-46fa-989f-fd6bb4e79769 · outbound

This paper cites Merging models with fisher-weighted averaging.Advances in Neural Information Processing Systems, 35:17703–17716, 2022.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model Merging models with fisher-weighted averaging.Advances in Neural Information Processing Systems, 35:17703–17716, 2022

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:13.853463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:13.853463Z digest=sha256:4bd3100cedb246f104b199c5f9cf9ba168cbe3baa6451eb7665f243f11c88475

Observation a3ac23ad-a909-40b9-81fe-b2e8d0973b32 · outbound

This paper cites Using an llm to help with code understanding.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model Using an llm to help with code understanding

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:13.950462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:13.950462Z digest=sha256:ddbe7d43fde049a32dba6cec6aa674c41847287cc872aba276b194718c815767

Observation 6c60b368-3315-473a-bc21-6a6e4f930483 · outbound

This paper cites The role of paralinguistic cues in social life.ANALYSIS OF MODERN SCIENCE AND INNOVATION, 1(2):215–218, 2024.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model The role of paralinguistic cues in social life.ANALYSIS OF MODERN SCIENCE AND INNOVATION, 1(2):215–218, 2024

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:03:18.931107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:03:14.045361Z digest=sha256:43fbbb8d0e44fc9b239507a39e127c995bbaceb20f7d537a46bfe37fe896eed2

Observation 2905d971-6f5a-4ca6-b4ba-96df5210b1ce · outbound

This paper cites A survey on speech large language models.arXiv preprint arXiv:2410.18908, 2024.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model A survey on speech large language models.arXiv preprint arXiv:2410.18908, 2024

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:14.141710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:14.141710Z digest=sha256:8f8a42bb2059ea1337f6cf70eb39dc45e27f93f6a1d9f268e5f9f812f487cf57

Observation cd1534b0-420e-4502-8a87-676fc3828313 · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.Advances in Neural Information Processing Systems, 36:53728–53741, 2023.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model Direct preference optimization: Your language model is secretly a reward model.Advances in Neural Information Processing Systems, 36:53728–53741, 2023

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:14.234520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:14.234520Z digest=sha256:c3241539283895d79fd0b8ef996a6c9b899896171ee324d315debd3612b5e3b8

Observation 7f2b0cb2-a3ab-483b-adf7-e8e47cc6535a · outbound

This paper cites AudioPaLM: A Large Language Model That Can Speak and Listen.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model AudioPaLM: A Large Language Model That Can Speak and Listen

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:14.333366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:14.333366Z digest=sha256:43a1f052f5422e463f3ea21ddbe7ff80ccf2da369a3ed78f650375f59176bd4b

Observation ddd32cd2-70b9-49a6-85db-d6918228071c · outbound

This paper cites How to debug code with github copilot.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model How to debug code with github copilot

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:03:18.762496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:03:14.429648Z digest=sha256:b19fba98cb24f63d8c6a47f390ab31b5d980b293d90534018339bcdd78ec9d9e

Observation 0450f964-3514-4b32-b6d4-1193c5f26a67 · outbound

This paper cites Paralinguistics in speech and language—state-of-the-art and the challenge.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model Paralinguistics in speech and language—state-of-the-art and the challenge

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:03:18.596275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:03:14.503974Z digest=sha256:fef3f118be2e38184ea77547207ac1e51239ebe909eb4594ec72c23c196ccd75

Observation 1f9684bb-3fd8-44ed-ba66-efc89785ae4c · outbound

This paper cites HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in Hugging Face.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in Hugging Face

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:14.618551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:14.618551Z digest=sha256:07a4b38d948cc42f0bbad100b2e892ebf4b19c602b5443bde01b928c0fe818d6

Observation 2c510c4f-08a8-4eb0-a358-fcd62f230dd2 · outbound

This paper cites Stepeval-audio-360.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model Stepeval-audio-360

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:03:18.443217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:03:14.752046Z digest=sha256:cf43f653efeeb968bcce9d4026e11a21b31cde8666b9c4297abf0bf965b4ff95

Observation 95470915-60f4-4c1e-8152-ab198d7116a5 · outbound

This paper cites SALMONN: Towards Generic Hearing Abilities for Large Language Models.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model SALMONN: Towards Generic Hearing Abilities for Large Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:14.876677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:14.876677Z digest=sha256:297c77a06c8ba89fa56cadf47db52d72ec04ac8eea5579f23abde0d710fd1128

Observation 81d875f8-529e-42f4-811b-7d2830f9f202 · outbound

This paper cites Preference alignment improves language model-based tts.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model Preference alignment improves language model-based tts

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:03:18.297584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:03:14.945288Z digest=sha256:428a76677b277d51a98529828c4cbbecc23fe6d58763c280c79e532d68285236

Observation d1195625-63a9-4c2a-b0c9-f18331437a73 · outbound

This paper cites Lami: Large language models for multi-modal human-robot interaction.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model Lami: Large language models for multi-modal human-robot interaction

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:03:18.113523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:03:15.071882Z digest=sha256:c2fe2e3cc8b79cdc2f75c4083696b25f194cb489d0b5eba942104ed8599a5781

Observation 907f9a29-8e0a-4a27-99f6-cd12850217ed · outbound

This paper cites Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:15.171753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:15.171753Z digest=sha256:019e21582435be47acd1960117d650f63d10fe10fb22ce2a1bd39867cedf6d73

Observation 41321313-bb77-41b8-856f-84ad7fcbe144 · outbound

This paper cites Reinforcement Learning for LLM Post-Training: A Survey.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model Reinforcement Learning for LLM Post-Training: A Survey

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:15.250044Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:15.250044Z digest=sha256:1c33e15a2c5b0c139ce72dd5beb96101fb8d81a553dcfaa263e4d087acaea126

Observation 66595b68-8294-47cb-94a7-89faf54611fd · outbound

This paper cites Model soups: aver- aging weights of multiple fine-tuned models improves accuracy without increasing inference time.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model Model soups: aver- aging weights of multiple fine-tuned models improves accuracy without increasing inference time

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:15.331162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:15.331162Z digest=sha256:c5d2ff58e81c537a1e41c595807a929890278e62fc9c6a2a00262e219291e920

Observation 962da2ef-045a-4b6e-b459-fe49ff99f826 · outbound

This paper cites Codec-SUPERB: An In-Depth Analysis of Sound Codec Models.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model Codec-SUPERB: An In-Depth Analysis of Sound Codec Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:15.425974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:15.425974Z digest=sha256:b05e58d9114b0d32a5b32568faf8f1cf39e0a9b9bfa7da729528921a5e6b24a1

Observation 07125a7f-1975-4d32-a803-21d414475f7b · outbound

This paper cites When search engine services meet large language models: visions and challenges.IEEE Transactions on Services Computing, 2024.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model When search engine services meet large language models: visions and challenges.IEEE Transactions on Services Computing, 2024

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:03:17.895203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:03:15.557323Z digest=sha256:d57f1a78aacffc9235328cf8de5518a810fc0dbf7b8af07047cbcdda94c6e149

Observation 7f1bb81c-d213-49f6-b752-269636b914d0 · outbound

This paper cites Qwen2.5-Omni Technical Report.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model Qwen2.5-Omni Technical Report

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:15.651302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:15.651302Z digest=sha256:457b85a5352663d73ed41a469206e0d1dfa5ed30358552b9efb3d197a7fd1d9b

Observation bd288521-e2f9-45f7-b906-647636d89695 · outbound

This paper cites Uniaudio: Towards universal audio generation with large language models.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model Uniaudio: Towards universal audio generation with large language models

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:03:17.676913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:03:15.749688Z digest=sha256:5d6621622899cdd655f5d182e21fe9b74077c4915e84810ba3f99fd72b1112e5

Observation 5eade042-b8c6-4f85-aa5f-faaa4851b64b · outbound

This paper cites Soundstream: An end-to-end neural audio codec.IEEE/ACM Transactions on Audio, Speech, and Language Processing, 30:495–507, 2021.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model Soundstream: An end-to-end neural audio codec.IEEE/ACM Transactions on Audio, Speech, and Language Processing, 30:495–507, 2021

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:15.837924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:15.837924Z digest=sha256:d4511a146b0244fae72a763d808ced2763b111232f1c4c2dd5961567c1271180

Observation 6364a05c-56a0-4e4d-a795-8ba0b4bfed1d · outbound

This paper cites GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:15.913730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:15.913730Z digest=sha256:ad63c53d0c452a7a2eaa66e3fef61a5b55c669e155c595cf44a1e179d0302aae

Observation fb0bdad5-184f-4cb4-a36d-75ce3e54ac98 · outbound

This paper cites Root mean square layer normalization.Advances in Neural Information Processing Systems, 32, 2019.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model Root mean square layer normalization.Advances in Neural Information Processing Systems, 32, 2019

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:15.976166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:15.976166Z digest=sha256:7fad08aaa3cd922f3995b9d35f4adeb0a3f3dffe71cf5c2d4068f1283d03723a

Observation 0d49440e-e10f-4e6a-9510-fbd1db9a59de · outbound

This paper cites SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:16.062120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:16.062120Z digest=sha256:bd41861bade215f606d5c663f9e8b906fa54e80f247307514bbaf7464043b2cf

Observation 359d011e-1276-4af3-b3e7-480e3cd1ae5c · outbound

This paper cites SpeechAlign: Aligning Speech Generation to Human Preferences.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model SpeechAlign: Aligning Speech Generation to Human Preferences

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:16.158400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:16.158400Z digest=sha256:fdb6614cd44190796cb63f496daeddab942936897ac3914d0c7e13f35fd96b26

Observation 2e36a9ec-f727-4c09-a3e8-6fcb7d392395 · outbound

This paper cites Vistorybench: Comprehensive benchmark suite for story visualiza- tion.arXiv preprint arXiv:2505.24862, 2025.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model Vistorybench: Comprehensive benchmark suite for story visualiza- tion.arXiv preprint arXiv:2505.24862, 2025

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:16.265675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:16.265675Z digest=sha256:251b58fa22472f34bb5f43b0ac91747c12ad980f1bd9e7ae1de82110485279a2

Observation 4e967e3e-1770-489c-9c03-e9a1b3228968 · outbound

This paper cites Toolqa: A dataset for llm question answering with external tools.Advances in Neural Information Processing Systems, 36:50117–50143, 2023.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model Toolqa: A dataset for llm question answering with external tools.Advances in Neural Information Processing Systems, 36:50117–50143, 2023

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:16.429765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:16.429765Z digest=sha256:c86d2d8221ff4ab90da1adc7d12f06d280b90eb41feab7c5b90da3572d96a730

Observation a601aaa1-bf4b-4671-b545-3f298fe43b31 · outbound

This paper cites Pre-trained language model based ranking in baidu search.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model Pre-trained language model based ranking in baidu search

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:03:17.385454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:03:16.518476Z digest=sha256:60e61f17babff9b68f76bc572ed32ecc36d1eda67a822da6cfe853372fbc4370

Pith citing papers

Observation 77ab2a72-895c-410e-9eb3-653860b0be7d · inbound

Step-Audio 2 Technical Report cites this paper.

Step-Audio 2 Technical Report Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-16T05:59:51.136199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-16T05:59:50.900436Z digest=sha256:0e8e4f3dcd46a2676b9fb90c5d76c59c95094b9dba78d4f72edcd2756c5535d4

Observation ac1ddd35-444d-4687-88c5-1ac7ceb4cc4e · inbound

WavAlign: Enhancing Intelligence and Expressiveness in Spoken Dialogue Models via Adaptive Hybrid Post-Training cites this paper.

WavAlign: Enhancing Intelligence and Expressiveness in Spoken Dialogue Models via Adaptive Hybrid Post-Training Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T11:10:09.183579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T11:06:07.355633Z digest=sha256:786ff08af06f04fb42b597d5183df903d7553b9a05647bac1df3126380b38bf9

Observation 86c944ae-9249-46eb-8c34-1251a9201072 · inbound

Aligning Language Models for Lyric-to-Melody Generation with Rule-Based Musical Constraints cites this paper.

Aligning Language Models for Lyric-to-Melody Generation with Rule-Based Musical Constraints Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-11T12:41:05.205344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T03:09:41.001660Z digest=sha256:d75e04e4ad7fbccf6b1a8ff0a9180a99011e61c8f82e042bcd192ec6a9498b01

Observation ab9d91c2-f249-4a97-9649-5b3a42591fff · inbound

VITA-QinYu: Expressive Spoken Language Model for Role-Playing and Singing cites this paper.

VITA-QinYu: Expressive Spoken Language Model for Role-Playing and Singing Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T04:50:56.014102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-11T01:03:09.942984Z digest=sha256:027a7c2050692c960f0e2e098e337fbbe0d60fed3056e561c6583c96bf2dcb86