Pith. sign in

Paper Citation Record · LEDGER

Vision-Integrated High-Quality Neural Speech Coding

As of 8 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 1 inbound Pith citation observation for arXiv:2505.23379.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.23379 v1

Coverage vector

measured 29 of 29 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:50:36.199609Z

measured 30 of 30 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:50:33.692524Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T12:50:36.461320Z

Reference resolution

29 of 29 outbound references displayed

  • verified exact0
  • verified fuzzy18
  • unresolved10
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4fbb2ba6-3512-4f00-9f03-ff89b0e0c602 · outbound

This paper cites Vision-Integrated High-Quality Neural Speech Coding.

Vision-Integrated High-Quality Neural Speech Coding Vision-Integrated High-Quality Neural Speech Coding

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T12:50:36.534950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:50:33.692524Z digest=sha256:e847eb84f242cb0162d2f25f7ddf300f9fd65a3e9fb4a2abb043d4cac6c3e144

Observation d6a41dc8-bbb7-4387-8cf5-a2c5edfb9303 · outbound

This paper cites Overview As shown in Figure 1, VNSC consists of a speech coding mod- ule, an image analysis-synthesis module and a feature fusion module.

Vision-Integrated High-Quality Neural Speech Coding Overview As shown in Figure 1, VNSC consists of a speech coding mod- ule, an image analysis-synthesis module and a feature fusion module

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:40.297099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:50:33.772725Z digest=sha256:c9d82608665869c8da9b7dab35ead3d9831cc6dc8f802c70d8c40ad2b7dc4197

Observation 25c20d15-83cc-4562-a600-fb73f8e5d5a4 · outbound

This paper cites an unresolved cited work.

Vision-Integrated High-Quality Neural Speech Coding Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:50:40.079579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:50:33.838446Z digest=sha256:0d8652025e3c2830896e2088c826d38b6c40b4f58fdcaea4486e57b5744a437e

Observation 98a283bd-3523-440a-b593-dedb784c6c8a · outbound

This paper cites an unresolved cited work.

Vision-Integrated High-Quality Neural Speech Coding Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:50:39.938610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:50:33.954738Z digest=sha256:efd2e0e148fb7f6ba455bd2fbd3204511311c74cf38200becb8b362991153eae

Observation 250044a9-f627-49dd-a427-5d32286a1307 · outbound

This paper cites VNSC is built upon the speech-modal MDCTCodec, with visual information extracted from lip images flowing into the speech coding process.

Vision-Integrated High-Quality Neural Speech Coding VNSC is built upon the speech-modal MDCTCodec, with visual information extracted from lip images flowing into the speech coding process

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:39.716689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:50:34.058108Z digest=sha256:6b9f35c68c137575687d6a325f8e569950a7fa276c662857e6de147c1cece123

Observation b63097c6-e681-49e1-87d4-7735c2a88809 · outbound

This paper cites Recommendation G.711: Pulse code modulation (PCM) of voice frequencies,.

Vision-Integrated High-Quality Neural Speech Coding Recommendation G.711: Pulse code modulation (PCM) of voice frequencies,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:39.517538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:50:34.158951Z digest=sha256:70515a1ea8295c341548f05c9ac25d4c9eb99898a3d5d56e0f245b5d5e84e906

Observation 09ef0972-0153-4e21-96e7-9c8649d87c8c · outbound

This paper cites Recommendation G.723: Speech coders for multimedia communications: dual-rate coder (5.3/6.3 kbps),.

Vision-Integrated High-Quality Neural Speech Coding Recommendation G.723: Speech coders for multimedia communications: dual-rate coder (5.3/6.3 kbps),

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:39.303607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:50:34.228261Z digest=sha256:c1a31a270b3a82164df505f4aa1e6ae3a7aca263b49b9823213e85b12311d09a

Observation 511bda03-b79d-4c90-acbf-19c4e09ddf7c · outbound

This paper cites Recommendation g.726: 40, 32, 24, and 16 kbps adap- tive differential pulse code modulation (ADPCM),.

Vision-Integrated High-Quality Neural Speech Coding Recommendation g.726: 40, 32, 24, and 16 kbps adap- tive differential pulse code modulation (ADPCM),

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:39.132658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:50:34.314412Z digest=sha256:3206ad8367393f5660a298531248c252529878237b0de2129c5a20387a825575

Observation deecb13e-b03e-47be-aa0d-0ddfc3c2cb34 · outbound

This paper cites Recommendation G.729: Coding of speech at 8 kbps using conjugate-structure algebraic-code-excited linear prediction (CS-ACELP),.

Vision-Integrated High-Quality Neural Speech Coding Recommendation G.729: Coding of speech at 8 kbps using conjugate-structure algebraic-code-excited linear prediction (CS-ACELP),

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:38.977607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:50:34.388617Z digest=sha256:0b2d88007b82ab2f4d2256f88a7462982d010a36f113e39177eb291acced0d85

Observation 31808cef-cef9-4d63-beca-2c361bb3c11d · outbound

This paper cites SoundStream: An end-to-end neural audio codec,.

Vision-Integrated High-Quality Neural Speech Coding SoundStream: An end-to-end neural audio codec,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T12:50:34.461142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:50:34.461142Z digest=sha256:eb1973d289abb209c43032cd7d31c3d16019c8c6f7a6d8c94df10955b9cb1c59

Observation 5091cf6c-f157-4820-8046-c2724f6058c4 · outbound

This paper cites A review of vector quantization tech- niques,.

Vision-Integrated High-Quality Neural Speech Coding A review of vector quantization tech- niques,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:38.841316Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:50:34.548372Z digest=sha256:df0af6c237fe197b91bf3e19d5c639e780974a017cc23d599d2aaeebccf42dd5

Observation 52c5127c-ac04-4d30-9de3-ce38584d7de3 · outbound

This paper cites High fidelity neural audio compression,.

Vision-Integrated High-Quality Neural Speech Coding High fidelity neural audio compression,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T12:50:34.631712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:50:34.631712Z digest=sha256:9f4d08f1a747d4078da4d3a2f69dd404821765869e5272910fb006d09eddd482

Observation cb3a9791-83df-42c5-b990-e4fc8f7d12cd · outbound

This paper cites HiFi-Codec: Group-residual Vector quantization for High Fidelity Audio Codec.

Vision-Integrated High-Quality Neural Speech Coding HiFi-Codec: Group-residual Vector quantization for High Fidelity Audio Codec

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T12:50:34.712516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:50:34.712516Z digest=sha256:89e6da3107d64c6de92e5fae54f5d6aa2bc6f4ffc5b15bc4131d0aa3bc8541a7

Observation 8551fd59-4311-4472-a9fa-9bd0a5ed5f46 · outbound

This paper cites APCodec: A neural audio codec with parallel amplitude and phase spec- trum encoding and decoding,.

Vision-Integrated High-Quality Neural Speech Coding APCodec: A neural audio codec with parallel amplitude and phase spec- trum encoding and decoding,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T12:50:34.799104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:50:34.799104Z digest=sha256:9f4fa59568485625aa003f38567df189a2fb9b53bd0d332b24d1c740e817b419

Observation aa39616b-e0e7-4ed1-827d-2321e6f97fc5 · outbound

This paper cites Mdctcodec: A lightweight mdct-based neural audio codec towards high sampling rate and low bitrate scenarios,.

Vision-Integrated High-Quality Neural Speech Coding Mdctcodec: A lightweight mdct-based neural audio codec towards high sampling rate and low bitrate scenarios,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:38.627209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:50:34.902677Z digest=sha256:e587b8c49cb3c725d01a92847e10e793300f1949fa5c8d6fff3c375473ac88aa

Observation f84ec758-6146-4190-922d-4311a50b54c1 · outbound

This paper cites DM-Codec: Distilling multimodal representations for speech tokenization,.

Vision-Integrated High-Quality Neural Speech Coding DM-Codec: Distilling multimodal representations for speech tokenization,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T12:50:35.013098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:50:35.013098Z digest=sha256:43d3fb9a684a9328fe137efe77cac3747382084414df2b66958f241576c9684a

Observation cf1d0463-7a56-4f46-87d5-20177b1622bf · outbound

This paper cites The conversation: Deep audio-visual speech enhancement,.

Vision-Integrated High-Quality Neural Speech Coding The conversation: Deep audio-visual speech enhancement,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:38.440004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:50:35.070455Z digest=sha256:2041f2d4a760bad33f1f6c539ce679bfc461b898215f2eee87bffa3f38c04768

Observation 2c7aecbf-bb65-4c15-890d-12399e068b1e · outbound

This paper cites Vsegan: Visual speech enhancement generative adversarial net- work,.

Vision-Integrated High-Quality Neural Speech Coding Vsegan: Visual speech enhancement generative adversarial net- work,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:38.250109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:50:35.180640Z digest=sha256:8f4e8c0172fcdcda743415749dbee033f507b1d1969fbf54586c176b4c2404b2

Observation 56d598f4-c43e-4b42-8d91-b09d2cb635e2 · outbound

This paper cites Incorporating ultra- sound tongue images for audio-visual speech enhancement,.

Vision-Integrated High-Quality Neural Speech Coding Incorporating ultra- sound tongue images for audio-visual speech enhancement,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:38.053040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:50:35.288448Z digest=sha256:9b84bb98aef17f9652f55c593876394a7f8193a870f7b8f8fbc49026c765d83f

Observation a31dba6f-16c4-422e-b95e-deba50925826 · outbound

This paper cites Improving visual speech enhancement network by learning audio-visual affinity with multi-head attention,.

Vision-Integrated High-Quality Neural Speech Coding Improving visual speech enhancement network by learning audio-visual affinity with multi-head attention,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:37.910741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:50:35.366485Z digest=sha256:14fb19199e95774e8acbbfded38f93d512c36efad9d7f022659a9dab339a9019

Observation fd840dd6-135b-4e53-a130-f89dff2afe3d · outbound

This paper cites TaL: a synchronised multi-speaker corpus of ultrasound tongue imaging, audio, and lip videos,.

Vision-Integrated High-Quality Neural Speech Coding TaL: a synchronised multi-speaker corpus of ultrasound tongue imaging, audio, and lip videos,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:37.731528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:50:35.413134Z digest=sha256:2a22091b2460427e5e966772c0910d649a79b015a0744433ff4c1a0eb76b6d31

Observation 5e24414c-2326-4c23-9d60-a652e68831ae · outbound

This paper cites Deep audio-visual speech recognition,.

Vision-Integrated High-Quality Neural Speech Coding Deep audio-visual speech recognition,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:37.563871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:50:35.521183Z digest=sha256:e17197f0ef86c61166ccb2910606a1e2ad728fd811bd20fa90aa9202895315f6

Observation da2cc957-4803-4b3f-b84b-4bd11fc504d5 · outbound

This paper cites ConvNeXt v2: Co-designing and scaling convnets with masked autoencoders,.

Vision-Integrated High-Quality Neural Speech Coding ConvNeXt v2: Co-designing and scaling convnets with masked autoencoders,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T12:50:35.615922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:50:35.615922Z digest=sha256:02cdd05bc1e9206847161c84175529bb5b2b8bec1736cc52bd54088f6c3ad73b

Observation 1ae59075-736a-471a-88ad-726a3c42d508 · outbound

This paper cites Gaussian Error Linear Units (GELUs).

Vision-Integrated High-Quality Neural Speech Coding Gaussian Error Linear Units (GELUs)

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T12:50:35.712227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:50:35.712227Z digest=sha256:f0843a981248d63adb78679785abf41c161a5def0a8be0b8360d857f55a585eb

Observation 53a7abad-c408-4b81-b797-fbde56d063a0 · outbound

This paper cites Converting video formats with ffmpeg,.

Vision-Integrated High-Quality Neural Speech Coding Converting video formats with ffmpeg,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:37.330103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:50:35.835330Z digest=sha256:c5deec896268496d4027dce4fb32c9fbf7150609c2b648a4035d6413460f1447

Observation b4f915c7-eaa2-4678-bab0-e134a6df0b8c · outbound

This paper cites Decoupled weight decay regulariza- tion,.

Vision-Integrated High-Quality Neural Speech Coding Decoupled weight decay regulariza- tion,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:37.107644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:50:35.961443Z digest=sha256:69cf5bfa3c6965562416c1ec2c559f7e69bc5bc6444567eefe9aad4f38345f05

Observation 2b135383-6ef9-4d95-9d66-31f8f24ec55b · outbound

This paper cites P. 862.2: Wideband extension to recom- mendation P. 862 for the assessment of wideband telephone networks and speech codecs,.

Vision-Integrated High-Quality Neural Speech Coding P. 862.2: Wideband extension to recom- mendation P. 862 for the assessment of wideband telephone networks and speech codecs,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:36.921223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:50:36.014763Z digest=sha256:09a8ac233ff102cf7bf90894f5c3e7879c6645b466be6141fdd26790943c868f

Observation 42100a44-3988-456e-b4d4-04bc84b8c6f0 · outbound

This paper cites Evaluation of objective measures for speech enhancement,.

Vision-Integrated High-Quality Neural Speech Coding Evaluation of objective measures for speech enhancement,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:36.696296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:50:36.121916Z digest=sha256:1b653920eb3645414bbd8a0b64b86f9e2928c39fba9fde0d0d9d552fd06178e5

Observation 66c807d6-df42-4ff4-870e-d991b8c967a8 · outbound

This paper cites A short- time objective intelligibility measure for time-frequency weighted noisy speech,.

Vision-Integrated High-Quality Neural Speech Coding A short- time objective intelligibility measure for time-frequency weighted noisy speech,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T12:50:36.199609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:50:36.199609Z digest=sha256:92e303e06c1c3f2f38f367edb7898c6fafc1983e42d0b9b1c1e81f9600d7621d

Pith citing papers

Observation 4fbb2ba6-3512-4f00-9f03-ff89b0e0c602 · inbound

Vision-Integrated High-Quality Neural Speech Coding cites this paper.

Vision-Integrated High-Quality Neural Speech Coding Vision-Integrated High-Quality Neural Speech Coding

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T12:50:36.534950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:50:33.692524Z digest=sha256:e847eb84f242cb0162d2f25f7ddf300f9fd65a3e9fb4a2abb043d4cac6c3e144