Pith. sign in

Paper Citation Record · LEDGER

Zipformer: A faster and better encoder for automatic speech recognition

As of 17 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 12 inbound Pith citation observations for arXiv:2310.11230.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2310.11230 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 12 of 12 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 12 of 12 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T19:58:55.031880Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T05:56:39.900954Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 28b369b5-17ba-4972-859f-618ccaef638b · inbound

Towards Maximum Likelihood Training for Transducer-based Streaming Speech Recognition cites this paper.

Towards Maximum Likelihood Training for Transducer-based Streaming Speech Recognition Zipformer: A faster and better encoder for automatic speech recognition

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T12:13:11.925137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:13:11.925137Z digest=sha256:def43d246dd5f8661ce7c3713f89da63f1e6649cd93ff3874fe458aa17b9761b

Observation 2ddd568e-09ad-4df4-bd24-95be75ee1743 · inbound

Pureformer-VC: Non-parallel Voice Conversion with Pure Stylized Transformer Blocks and Triplet Discriminative Training cites this paper.

Pureformer-VC: Non-parallel Voice Conversion with Pure Stylized Transformer Blocks and Triplet Discriminative Training Zipformer: A faster and better encoder for automatic speech recognition

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T05:20:56.462348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:20:56.462348Z digest=sha256:ac71c6e859e69c26e88d1fe21c34f7b4d4573bd6c805abf00900a0dbb3952f55

Observation d169ada8-5607-4285-9f03-1827045ef359 · inbound

Unifying Streaming and Non-streaming Zipformer-based ASR cites this paper.

Unifying Streaming and Non-streaming Zipformer-based ASR Zipformer: A faster and better encoder for automatic speech recognition

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T19:58:55.031880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:58:55.031880Z digest=sha256:7c36467cc3572a92aded274aa3353a5c802b33ce1bfb59570798418eae54efad

Observation 5835c07c-950d-4118-957f-594a85fa128b · inbound

Robust and Efficient Autoregressive Speech Synthesis with Dynamic Chunk-wise Prediction Policy cites this paper.

Robust and Efficient Autoregressive Speech Synthesis with Dynamic Chunk-wise Prediction Policy Zipformer: A faster and better encoder for automatic speech recognition

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T22:19:03.875806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:19:03.875806Z digest=sha256:065998077f6654622563cf5ba40dd0614c974bc97f52ced140476e4b37dcc8d5

Observation a022cac8-8adc-4ed1-b89f-2c25917be940 · inbound

BoSS: Beyond-Semantic Speech cites this paper.

BoSS: Beyond-Semantic Speech Zipformer: A faster and better encoder for automatic speech recognition

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T14:50:04.324971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:50:04.324971Z digest=sha256:5156707ca99f8cf8ea2a08599dfa436f8180c1d582213312a5d7ac4a7bef84d6

Observation a8dc773a-6858-428f-aa15-02d2f48e913f · inbound

Towards Hallucination-Free Music: A Reinforcement Learning Preference Optimization Framework for Reliable Song Generation cites this paper.

Towards Hallucination-Free Music: A Reinforcement Learning Preference Optimization Framework for Reliable Song Generation Zipformer: A faster and better encoder for automatic speech recognition

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-05T23:43:05.627597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:43:05.627597Z digest=sha256:9b513fb3366e65d7e281c247ea625903f376bad482dd34a810dd39488bb6790d

Observation 36270e26-3b47-4623-b5ba-0d932e9939a4 · inbound

XEmoRAG: Cross-Lingual Emotion Transfer with Controllable Intensity Using Retrieval-Augmented Generation cites this paper.

XEmoRAG: Cross-Lingual Emotion Transfer with Controllable Intensity Using Retrieval-Augmented Generation Zipformer: A faster and better encoder for automatic speech recognition

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-05T22:17:38.566540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:17:38.566540Z digest=sha256:62a153439a87672cabb521565e85fc29228014bfa82e5db16473e0444477d652

Observation e171fbec-8b85-4df6-ac32-35cbfffa24d1 · inbound

Exploring Cross-Utterance Speech Contexts for Conformer-Transducer Speech Recognition Systems cites this paper.

Exploring Cross-Utterance Speech Contexts for Conformer-Transducer Speech Recognition Systems Zipformer: A faster and better encoder for automatic speech recognition

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-05T20:31:45.751156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:31:45.751156Z digest=sha256:d2620d55a928b8c20b47eeb141de12523f1bdc56e724914d51514f12a4326d78

Observation c6f16ced-7f86-4346-bbcb-62104c5b69fa · inbound

UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models cites this paper.

UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models Zipformer: A faster and better encoder for automatic speech recognition

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-04T11:29:37.020821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T11:29:37.020821Z digest=sha256:37c76ffbfa1fea3f1194ddb6291463c06b1049afdbbbfbc81ce0c42b167e9dd5

Observation 4b380ae4-071b-42b5-9449-f9eb6e089af6 · inbound

Raon-Speech Technical Report cites this paper.

Raon-Speech Technical Report Zipformer: A faster and better encoder for automatic speech recognition

Reference 34

Resolution
unresolved
no resolver link, observed 2026-07-13T08:22:44.347941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T08:22:44.347941Z digest=sha256:42ab49ae96e809806ea9b7d5446b23b908817c623b6d4cb8d4b04f78f4e17cdb

Observation 18c5fc83-4590-4316-8f63-6112e91f77e4 · inbound

SketchSong: Hierarchical Song Generation with Sketch Planning and Fine-Grained Multi-Track Modeling cites this paper.

SketchSong: Hierarchical Song Generation with Sketch Planning and Fine-Grained Multi-Track Modeling Zipformer: A faster and better encoder for automatic speech recognition

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-07-02T04:46:38.053008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-28T08:53:17.009342Z digest=sha256:74229b35acc70aa283824c6bdacee77b5361f27ec76634b89d8be5f53fff865b

Observation b2610e2d-264d-4edb-aa02-8b362b035575 · inbound

From Objectives to Applications: Aligning Architectural Biases in Audio Self-Supervised Learning cites this paper.

From Objectives to Applications: Aligning Architectural Biases in Audio Self-Supervised Learning Zipformer: A faster and better encoder for automatic speech recognition

Reference 76

Resolution
verified exact
arxiv_id, observed 2026-07-02T05:56:39.902887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-07-02T05:52:55.818877Z digest=sha256:83a273f4efca697103ba21579485d0ea17cf693ce4fd5f6edf1f601376cd41df