Pith. sign in

Paper Citation Record · LEDGER

StepAudio 2.5 Technical Report

As of 10 August 2026, this Paper Citation Record lists 39 of 39 outbound references and 4 inbound Pith citation observations for arXiv:2605.23463.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.23463 v1

Coverage vector

measured 39 of 39 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-25T02:52:22.610397Z

measured 43 of 43 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T18:13:04.179577Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-03T17:18:44.075741Z

Reference resolution

39 of 39 outbound references displayed

  • verified exact20
  • verified fuzzy18
  • unresolved1
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 086e75aa-a6ef-417f-b944-038c36ad785f · outbound

This paper cites Connectionist temporal classification.

StepAudio 2.5 Technical Report Connectionist temporal classification

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T02:56:35.361557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T02:52:22.610397Z digest=sha256:b6dbb9ef0effe90d6dfc9a0d550f59357f5780c2b969783f24786ce64b438c58

Observation 14ae607c-95e5-44d1-bdf8-4b5aacb3f70c · outbound

This paper cites Sequence Transduction with Recurrent Neural Networks.

StepAudio 2.5 Technical Report Sequence Transduction with Recurrent Neural Networks

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-25T02:55:16.521251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T02:52:22.610397Z digest=sha256:fb7ac023553fd00fc9155b3aa0cfb76d78b2052b929c5c2f7e86cc803efb89fd

Observation 6eab5b32-aab8-4e2a-a3fe-7d9fdbfc6b2a · outbound

This paper cites Listen, Attend and Spell.

StepAudio 2.5 Technical Report Listen, Attend and Spell

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-25T02:55:16.509973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T02:52:22.610397Z digest=sha256:0e6dfd367c8942d6f68a9332fc6b42505e14603f8d93ada41029a0f9cfc98702

Observation 5eeac4c3-be4e-46d5-b52c-f412f6d551a2 · outbound

This paper cites Robust speech recognition via large-scale weak supervision.

StepAudio 2.5 Technical Report Robust speech recognition via large-scale weak supervision

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T02:56:35.353701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T02:52:22.610397Z digest=sha256:7147695d0124bb809bdc7e081a2b4bf66f2772df6395a3cd292c99fe5bf5fca6

Observation df029add-8e4b-4238-ab59-6d9334d01e84 · outbound

This paper cites VIBEVOICE-ASR technical report.

StepAudio 2.5 Technical Report VIBEVOICE-ASR technical report

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-25T02:55:16.526933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T02:52:22.610397Z digest=sha256:5323673214f6f9894e44343000366eb428f485e0bc36ecebe868d98e909a5fea

Observation 5e20d531-5ce5-478e-afeb-3a15edf3fa37 · outbound

This paper cites Fun-ASR technical report.

StepAudio 2.5 Technical Report Fun-ASR technical report

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-25T02:55:16.532217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T02:52:22.610397Z digest=sha256:2d730dfdd795e6d3e4a4d692fefbfa3101a8636d4b42b59097b5aad74203b953

Observation 91001d88-a7f2-4930-8541-dc76486a95f5 · outbound

This paper cites Seed-ASR: Understanding Diverse Speech and Contexts with LLM-based Speech Recognition.

StepAudio 2.5 Technical Report Seed-ASR: Understanding Diverse Speech and Contexts with LLM-based Speech Recognition

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-25T02:55:16.515698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T02:52:22.610397Z digest=sha256:52cedcd61982aac0f23630136af7fc20e4fcd3ec4f8282c9a0faffbd2e3a2050

Observation 9fd13268-45a3-43f5-a739-bfc06b0f025e · outbound

This paper cites Qwen3-ASR Technical Report.

StepAudio 2.5 Technical Report Qwen3-ASR Technical Report

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-25T02:55:16.483222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T02:52:22.610397Z digest=sha256:686be71945020a62c186e906e49e3bc9a1a87c3e3d48977fd0629bf7339a6844

Observation ae35f11f-b45b-4570-a385-d37c6d821ce6 · outbound

This paper cites Step-Audio 2 Technical Report.

StepAudio 2.5 Technical Report Step-Audio 2 Technical Report

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-25T02:55:16.487852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T02:52:22.610397Z digest=sha256:780b8f7cf3b5c3a6b9b343020c88457b25eb134dd016dbb6744b405e2ff0a956

Observation 713f12d6-bd4a-407b-8965-ba391750f8e0 · outbound

This paper cites Qwen3-Omni Technical Report.

StepAudio 2.5 Technical Report Qwen3-Omni Technical Report

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-25T02:55:16.441821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T02:52:22.610397Z digest=sha256:c81d298b63186317591582e636ba7fe9f773758c5dc81d9d918903cb4cb6287e

Observation 0953c54f-6449-4679-b7a8-a5382add346c · outbound

This paper cites Boosting Omni-Modal Language Models: Staged Post-Training with Visually Debiased Evaluation.

StepAudio 2.5 Technical Report Boosting Omni-Modal Language Models: Staged Post-Training with Visually Debiased Evaluation

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-25T02:55:16.477459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T02:52:22.610397Z digest=sha256:d1ed9b77c5e911d05ba9fbd9906eeaf70aade253638dde8b6641fa2684f6ffec

Observation fc8064bb-97ff-403c-9dfd-eedabf2c2453 · outbound

This paper cites Salmonn: Towards generic hearing abilities for large language models.

StepAudio 2.5 Technical Report Salmonn: Towards generic hearing abilities for large language models

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T02:56:35.371149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T02:52:22.610397Z digest=sha256:34455568175e4be332daed91dc98edb1c1ba173662c8577b48c512fa79e4a31f

Observation 57515c3d-9c92-4e8e-9189-36d3eb900b8a · outbound

This paper cites Audiolm: a language modeling approach to audio generation.IEEE/ACM transactions on audio, speech, and language processing, 31:2523–2533.

StepAudio 2.5 Technical Report Audiolm: a language modeling approach to audio generation.IEEE/ACM transactions on audio, speech, and language processing, 31:2523–2533

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T02:56:35.374738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T02:52:22.610397Z digest=sha256:afa8ea6d4d2d2ae067ff5a686c8be30bfb6dd14948f505f46ffc3c244dd150b1

Observation 350987e5-fa75-45d3-8571-0bcdb7a12996 · outbound

This paper cites Recent advances in speech language models: A survey.

StepAudio 2.5 Technical Report Recent advances in speech language models: A survey

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T02:56:35.379842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T02:52:22.610397Z digest=sha256:8a593a02cdaff1abefb3040407dd934465ea7bb99c8a53188ee8a0ed628e9ebb

Observation c59d0276-31d1-4f4f-b648-9517b90875b5 · outbound

This paper cites Paralinguistics-aware speech-empowered large language models for natural conversation.Advances in Neural Information Processing Systems, 37:131072–131103.

StepAudio 2.5 Technical Report Paralinguistics-aware speech-empowered large language models for natural conversation.Advances in Neural Information Processing Systems, 37:131072–131103

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T02:56:35.383483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T02:52:22.610397Z digest=sha256:e64261414e1dbac23a0bfa269be0d592408a36ed5c0e037c802c2a5145c03501

Observation 0caf796e-e190-4ba0-b488-fdd8162c462c · outbound

This paper cites Freeze-omni: A smart and low latency speech-to-speech dialogue model with frozen LLM.

StepAudio 2.5 Technical Report Freeze-omni: A smart and low latency speech-to-speech dialogue model with frozen LLM

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T02:56:35.389815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T02:52:22.610397Z digest=sha256:2b7da47e607c46fd9a1ef1df84ee9448d65eebf124c4bf95fee1641a8222f117

Observation 2c452943-18a0-4a86-b196-e4b4c9f47119 · outbound

This paper cites Depflow: Disentangled speech generation to mitigate semantic bias in depression detection.

StepAudio 2.5 Technical Report Depflow: Disentangled speech generation to mitigate semantic bias in depression detection

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-25T02:55:16.431713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T02:52:22.610397Z digest=sha256:d4a939e815f1049b6588c6eb10e8417612cc15772ecc17217ebe879b558c7a52

Observation 9b34497f-a2a0-43e7-a002-f32ab6dbe7ee · outbound

This paper cites A new approach to extract fetal electrocardiogram using affine combination of adaptive filters.

StepAudio 2.5 Technical Report A new approach to extract fetal electrocardiogram using affine combination of adaptive filters

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T02:56:35.367549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T02:52:22.610397Z digest=sha256:4f12a43b31764e2171718b035b6cee3fc89fafe90d1537d949f4d5e83b76e244

Observation e2be10c6-91b2-40e5-abf2-c6973cc89d8b · outbound

This paper cites Multi-bench: A multi-turn interactive benchmark for assessing emotional intelligence ability of spoken dialogue models.

StepAudio 2.5 Technical Report Multi-bench: A multi-turn interactive benchmark for assessing emotional intelligence ability of spoken dialogue models

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-25T02:55:16.436892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T02:52:22.610397Z digest=sha256:2f0e668f677a8ca625d0418e97601b10f6f3ee19087ce0f993926c671e62726e

Observation be94f0cd-1a89-49c5-9cc2-8a8b361e35e7 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

StepAudio 2.5 Technical Report Gemini: A Family of Highly Capable Multimodal Models

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-25T02:55:16.463143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T02:52:22.610397Z digest=sha256:05203dd2e11c4d4ad95b5f59769a59a5032a2dbf3449116776e3d130f29925b3

Observation a0b65575-9b51-44bd-9706-5a1487dd1153 · outbound

This paper cites Mind-Paced Speaking: A Dual-Brain Approach to Real-Time Reasoning in Spoken Language Models.

StepAudio 2.5 Technical Report Mind-Paced Speaking: A Dual-Brain Approach to Real-Time Reasoning in Spoken Language Models

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-05-25T02:55:16.472393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T02:52:22.610397Z digest=sha256:c6d7545d69029ffee54bb6e7ac7f4da19b066eeaa0c4cd8325e868fd78cca5ae

Observation 3f84023c-a4ee-4d5f-bc42-95785eff768a · outbound

This paper cites Chronological thinking in full-duplex spoken dialogue language models.

StepAudio 2.5 Technical Report Chronological thinking in full-duplex spoken dialogue language models

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-25T02:55:16.505381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T02:52:22.610397Z digest=sha256:6e47897b761249297d7491f4a2adc9a2f51cf109a6bd9e443132a7a9b3c33b1b

Observation 63082297-3c5c-4db9-aca2-327d8a2db705 · outbound

This paper cites Duplexsla: A full-duplex spoken language model with synchronized speech, language, and action.

StepAudio 2.5 Technical Report Duplexsla: A full-duplex spoken language model with synchronized speech, language, and action

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T02:56:35.386641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T02:52:22.610397Z digest=sha256:f23f44f68b4f472e7e0dfd59eddfbc31e5af8dd78f8866c09a9f16cbc3ee9858

Observation 20c44ea9-682f-4e47-ab7b-0581c94a5b9b · outbound

This paper cites Mamba in speech: Towards an alternative to self-attention.IEEE Transactions on Audio, Speech and Language Processing.

StepAudio 2.5 Technical Report Mamba in speech: Towards an alternative to self-attention.IEEE Transactions on Audio, Speech and Language Processing

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T02:56:35.392860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T02:52:22.610397Z digest=sha256:bd68060809b8da04cf26dd6a6eb8d4bf5e2685a5480412b939a8d423d81c5453

Observation df5d89ca-d3a6-4fe0-b631-566a1f8999a9 · outbound

This paper cites Code-switching speech recognition under the lens: Model-and data-centric perspectives.IEEE Transactions on Audio, Speech and Language Processing.

StepAudio 2.5 Technical Report Code-switching speech recognition under the lens: Model-and data-centric perspectives.IEEE Transactions on Audio, Speech and Language Processing

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T02:56:35.396848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T02:52:22.610397Z digest=sha256:dbad72f40366a266ae288385e39f6424464274c4564ecd6b5685f6f1ac009be6

Observation 413f7c81-0419-41ee-9e47-258ff55a93e6 · outbound

This paper cites Step-audio-r1 technical report.

StepAudio 2.5 Technical Report Step-audio-r1 technical report

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-25T02:55:16.500260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T02:52:22.610397Z digest=sha256:4517bbf31a9db252fe9af0218e96f8650f36bf274a3acf4f05299a517829fe7e

Observation 61a5a62a-d7e4-43fe-b7da-eb8aaabcd02b · outbound

This paper cites Step-Audio-R1.5 Technical Report.

StepAudio 2.5 Technical Report Step-Audio-R1.5 Technical Report

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-05-25T02:55:16.494617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T02:52:22.610397Z digest=sha256:2936f7913e75031e8cb6dadf07bb174e845e619ed04da26b286d518d6cf83a13

Observation 6bef68ea-45ae-41f8-adba-1faf1a74e168 · outbound

This paper cites Park, William Chan, Yu Zhang, et al.

StepAudio 2.5 Technical Report Park, William Chan, Yu Zhang, et al

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T02:56:35.402758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T02:52:22.610397Z digest=sha256:b37d0aa508b20a1d87b49f83cc72c17c55e6290252420657ef79aa7a5ce96c75

Observation 1a6e80ea-7904-4bf1-ba00-a51938a47ef3 · outbound

This paper cites an unresolved cited work.

StepAudio 2.5 Technical Report Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-05-25T02:56:35.406529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T02:52:22.610397Z digest=sha256:17eeaf888398311c4870ea2ef285cd0e7511b80379afb966081bfa3720810f67

Observation 697f37fc-63c2-408f-8eb4-c9cef2fdab7e · outbound

This paper cites AIShell-1: An open-source mandarin speech corpus and a speech recognition baseline.

StepAudio 2.5 Technical Report AIShell-1: An open-source mandarin speech corpus and a speech recognition baseline

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T02:56:35.409728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T02:52:22.610397Z digest=sha256:dfad1484b191519d25e4c465de5d3426e2e20002ec174472cb8c0a1048314f77

Observation 01cf541e-269d-45f9-9451-c73612401bd9 · outbound

This paper cites AISHELL-2: Transforming Mandarin ASR Research Into Industrial Scale.

StepAudio 2.5 Technical Report AISHELL-2: Transforming Mandarin ASR Research Into Industrial Scale

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-05-25T02:55:16.451910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T02:52:22.610397Z digest=sha256:1292c7224dbfa16cbb8285622d91a183bc5466efd17542a8515fdce01fccaa82

Observation 8d2d211f-feb6-4434-83a8-468b733a68fc · outbound

This paper cites WenetSpeech: A 10000+ hours multi-domain mandarin corpus for speech recognition.

StepAudio 2.5 Technical Report WenetSpeech: A 10000+ hours multi-domain mandarin corpus for speech recognition

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T02:56:35.422611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T02:52:22.610397Z digest=sha256:7179c4ea62fdd7987fecdebfb63f7cf850f24d0e3b0a6fab19e6f5f4c23280cb

Observation 49a49c25-9c27-466e-ae3e-3b83d2481bb9 · outbound

This paper cites FLEURS: Few-shot Learning Evaluation of Universal Representations of Speech.

StepAudio 2.5 Technical Report FLEURS: Few-shot Learning Evaluation of Universal Representations of Speech

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-25T02:55:16.447086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T02:52:22.610397Z digest=sha256:8614ac84d04a4bb0901e3a1ff17353b42c6e5d29d6a0d66d9b1520b7d10caa61

Observation 01041449-4711-47c4-8cb1-e0f0585030bc · outbound

This paper cites LibriSpeech: An ASR corpus based on public domain audio books.

StepAudio 2.5 Technical Report LibriSpeech: An ASR corpus based on public domain audio books

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T02:56:35.425885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T02:52:22.610397Z digest=sha256:74804319d639bf5d24f6430699d8545a43a5e003d9f2d18dbd1d6433abd0ece3

Observation 8f1435fe-f4a1-4c63-b748-925370e9a75e · outbound

This paper cites Common voice: A massively-multilingual speech corpus.

StepAudio 2.5 Technical Report Common voice: A massively-multilingual speech corpus

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T02:56:35.419215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T02:52:22.610397Z digest=sha256:e01ab02e08ff898ecfc7eea82ba03c36d3ae626bff25701c2867c3a91d374875

Observation 6c171711-18e3-4de3-9885-1f7b8eac9738 · outbound

This paper cites V oxpopuli-cleaned-aa: Cleaned ground truth transcripts for voxpopuli english test set.

StepAudio 2.5 Technical Report V oxpopuli-cleaned-aa: Cleaned ground truth transcripts for voxpopuli english test set

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T02:56:35.415860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T02:52:22.610397Z digest=sha256:a18153f55b552284712d294b32424b7913ae091dc75311fc4ed531cf8f48a223

Observation 9d1bfb95-8e08-47a5-b14c-f7e005da005c · outbound

This paper cites Earnings22-cleaned-aa: Cleaned ground truth transcripts for earnings22 english test set.

StepAudio 2.5 Technical Report Earnings22-cleaned-aa: Cleaned ground truth transcripts for earnings22 english test set

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T02:56:35.412796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T02:52:22.610397Z digest=sha256:e7e9c898e53cabfc8e7b1a35f0888e854a369e626ae9dcbe768bcb186bc9455a

Observation 4534e43b-4358-4eec-be62-10f5a597de92 · outbound

This paper cites Step-audio-editx technical report.

StepAudio 2.5 Technical Report Step-audio-editx technical report

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-25T02:55:16.458036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T02:52:22.610397Z digest=sha256:4987a70bbddd23a6d89904c86630fb0c419b12433d5741428077bd840e2eecbe

Observation ae24bc43-fe61-4d31-a7a5-2443a782c1a9 · outbound

This paper cites Proximal Policy Optimization Algorithms.

StepAudio 2.5 Technical Report Proximal Policy Optimization Algorithms

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-05-25T02:55:16.467568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T02:52:22.610397Z digest=sha256:4c13e4ac3dd0d7015c222b9730c8722397870613f9456e6cec28c692efc4c5cb

Pith citing papers

Observation 789faeb4-f885-46e4-a64f-704563c7b88e · inbound

Improving Code-Switching ASR with Code-Mixing Guided Synthetic Speech cites this paper.

Improving Code-Switching ASR with Code-Mixing Guided Synthetic Speech StepAudio 2.5 Technical Report

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-07-03T17:18:44.077038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T04:19:13.689689Z digest=sha256:1ad58b8931c763680364bc68c9efac91d8ae2395f7a19a74624f3e7a1032d9f8

Observation a05dba3a-3f4d-45fe-ad13-fa4c033e29f2 · inbound

ParaASR: Multi-Token Prediction for Fast and Long-Context LLM-Based Speech Recognition cites this paper.

ParaASR: Multi-Token Prediction for Fast and Long-Context LLM-Based Speech Recognition StepAudio 2.5 Technical Report

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-03T10:13:56.110991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:13:56.110991Z digest=sha256:682b0cb92719fb5720372acea7a819f95609509892dfcde34c9b523904716350

Observation e4b4e415-38b6-40d6-9aca-411f1eb58d31 · inbound

Breaking the Curse of Multilinguality in Many-to-Many Speech-to-Text Translation via a Resource-Aware Mixture of Speech Encoders cites this paper.

Breaking the Curse of Multilinguality in Many-to-Many Speech-to-Text Translation via a Resource-Aware Mixture of Speech Encoders StepAudio 2.5 Technical Report

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T21:28:33.190018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:28:33.190018Z digest=sha256:8644ee2f167812c0d5388d714767d50838b4eb82fd7495fc8f63c482ad426e75

Observation 090e4f5b-a5f6-48b6-8b13-4f2c233b7f3c · inbound

Breaking the Curse of Multilinguality in Many-to-Many Speech-to-Text Translation via a Resource-Aware Mixture of Speech Encoders cites this paper.

Breaking the Curse of Multilinguality in Many-to-Many Speech-to-Text Translation via a Resource-Aware Mixture of Speech Encoders StepAudio 2.5 Technical Report

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-08T18:13:04.179577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T18:13:04.179577Z digest=sha256:c2e21a06841712bc5160ba2211dbba45756bcb551b1c61916b234d9c2596b255