Pith. sign in

Paper Citation Record · LEDGER

Optimizing Speech Multi-View Feature Fusion through Conditional Computation

As of 19 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 0 inbound Pith citation observations for arXiv:2501.08057.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.08057 v1

Coverage vector

measured 31 of 31 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T20:35:13.120441Z

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

31 of 31 outbound references displayed

  • verified exact1
  • verified fuzzy17
  • unresolved13
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 88c42a83-1785-410c-8ab9-fb90274451a7 · outbound

This paper cites Introduction to digital speech processing,.

Optimizing Speech Multi-View Feature Fusion through Conditional Computation Introduction to digital speech processing,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:35:13.676357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T20:35:12.952741Z digest=sha256:12a710ca54103f487e4ea3360228856880fa66d0457518ef2cfadc85305386ad

Observation 1937764e-8328-45c1-90e1-9d91f823195f · outbound

This paper cites Learning robust features using deep learning for automatic seizure detection,.

Optimizing Speech Multi-View Feature Fusion through Conditional Computation Learning robust features using deep learning for automatic seizure detection,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:35:13.664268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T20:35:12.957265Z digest=sha256:67ca8f4ad047a63f2b753e0f3e552a204193276361de7c829e8c2fc91b6f1ef6

Observation de5af010-494a-43ef-a812-c3d97c274977 · outbound

This paper cites Listen, Attend and Spell.

Optimizing Speech Multi-View Feature Fusion through Conditional Computation Listen, Attend and Spell

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:12.962138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:12.962138Z digest=sha256:714fc1a17624e9eff3936190affc26219c5d105843e59f3ce8715c43a7e9ebfa

Observation dcc3f86b-bba6-49b1-bb64-2770c4ced7f2 · outbound

This paper cites wav2vec: Unsupervised Pre-training for Speech Recognition.

Optimizing Speech Multi-View Feature Fusion through Conditional Computation wav2vec: Unsupervised Pre-training for Speech Recognition

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:12.967080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:12.967080Z digest=sha256:148cbb94276eee7588f778857f66548145c84c9a4b6378fc66659df17b37352a

Observation 1be994b9-d8e9-43f4-af09-481ad06a238a · outbound

This paper cites Data2vec: A general framework for self-supervised learning in speech, vision and language,.

Optimizing Speech Multi-View Feature Fusion through Conditional Computation Data2vec: A general framework for self-supervised learning in speech, vision and language,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:35:13.652416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T20:35:12.972165Z digest=sha256:301cd9a9b03e0fe6d810b71a800ffc9aa3be01709fd7b2cb412b1f85b2710d01

Observation 134a1727-a30e-4c74-9574-32c86366f7ce · outbound

This paper cites wav2vec 2.0: A framework for self-supervised learning of speech representations,.

Optimizing Speech Multi-View Feature Fusion through Conditional Computation wav2vec 2.0: A framework for self-supervised learning of speech representations,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:35:13.640340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T20:35:12.978136Z digest=sha256:0fc7b52b2f418e9695c0d20287639218bab289a2c325b593580c7661c3de5f58

Observation 809aaec6-045b-47cc-a99c-3980a4e25ae2 · outbound

This paper cites Hubert: Self- supervised speech representation learning by masked prediction of hidden units,.

Optimizing Speech Multi-View Feature Fusion through Conditional Computation Hubert: Self- supervised speech representation learning by masked prediction of hidden units,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:35:13.628277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T20:35:12.983242Z digest=sha256:42b2e014f9ae6281ad0b18b6156e83b60d901669285e8d65a99ff8012951556e

Observation b4f68caa-92d2-4541-a491-63382220b722 · outbound

This paper cites Exploration on HuBERT with Multiple Resolu- tions,.

Optimizing Speech Multi-View Feature Fusion through Conditional Computation Exploration on HuBERT with Multiple Resolu- tions,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:35:13.616277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T20:35:12.988042Z digest=sha256:6300ee6f0f06e45c5fe79b96ab117e18e1835a89baf5e74e3208b87e49415f00

Observation e9db1f10-40da-4977-beff-caa965cdb38f · outbound

This paper cites Multi-resolution HuBERT: Multi-resolution Speech Self-Supervised Learning with Masked Unit Prediction.

Optimizing Speech Multi-View Feature Fusion through Conditional Computation Multi-resolution HuBERT: Multi-resolution Speech Self-Supervised Learning with Masked Unit Prediction

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:12.992523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:12.992523Z digest=sha256:d140d61463a0342056de8b5fd16fc971865bf714d3a2d39b815b64a4bc20d6e0

Observation 7b03d0e0-e681-4756-a5ba-ad0797c0c66a · outbound

This paper cites UnitSpeech: Speaker-adaptive Speech Synthesis with Untranscribed Data.

Optimizing Speech Multi-View Feature Fusion through Conditional Computation UnitSpeech: Speaker-adaptive Speech Synthesis with Untranscribed Data

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:12.996937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:12.996937Z digest=sha256:075f459517cc08274f28bc338f6fb42ba08a77b8d9a8cf75bdc46e840d00b118

Observation 79750105-313c-41cc-81e7-5be56738d061 · outbound

This paper cites SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities.

Optimizing Speech Multi-View Feature Fusion through Conditional Computation SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:13.001744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:13.001744Z digest=sha256:89758a7673dceca60e73a3bde61d5b10c5d7873dcf9edfbcdbde00b90350e30f

Observation 4ee0998a-0bc6-43dc-9c7c-0a006c6fc662 · outbound

This paper cites Multi- view information-bottleneck representation learning,.

Optimizing Speech Multi-View Feature Fusion through Conditional Computation Multi- view information-bottleneck representation learning,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:35:13.604733Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T20:35:13.007235Z digest=sha256:d2d604c9eb8231123041bed401c93e8785f45a8f92e819afbbd792339b821ae2

Observation 97adc650-762e-429e-830a-ce6f1a778a6f · outbound

This paper cites Deep multi-view learning methods: A review,.

Optimizing Speech Multi-View Feature Fusion through Conditional Computation Deep multi-view learning methods: A review,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:35:13.589761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T20:35:13.011623Z digest=sha256:50a089a777813eaba80dd607a1f20724a2dea66d6116a81efc00cb2fb7c98b7f

Observation 3cf63879-175c-4448-8c8d-431e7a867dbe · outbound

This paper cites The Platonic Representation Hypothesis.

Optimizing Speech Multi-View Feature Fusion through Conditional Computation The Platonic Representation Hypothesis

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:13.016293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:13.016293Z digest=sha256:e124413e62570884779a0a103df9ee6417a7c380b0289df833ce93f551005adb

Observation 80e901f7-d97f-4080-a2ec-78a4f9ec1370 · outbound

This paper cites Combining Spectral and Self-Supervised Features for Low Resource Speech Recognition and Translation,.

Optimizing Speech Multi-View Feature Fusion through Conditional Computation Combining Spectral and Self-Supervised Features for Low Resource Speech Recognition and Translation,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:35:13.574820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T20:35:13.023101Z digest=sha256:5cf88dbdae8337e6063e8861e5be9e20b8bfad408a5f755e3a5de0971943dc67

Observation 42892fa8-b830-4431-b653-a619d8cf2d38 · outbound

This paper cites Improving speech emotion recognition by fusing self-supervised learn- ing and spectral features via mixture of experts,.

Optimizing Speech Multi-View Feature Fusion through Conditional Computation Improving speech emotion recognition by fusing self-supervised learn- ing and spectral features via mixture of experts,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:35:13.559020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T20:35:13.028995Z digest=sha256:48755d7859717c5991474904f5d90f20af1878d9984625aed5f913d9cfdf2269

Observation 7db3994b-4a62-4b3c-a152-892cfc5cc83b · outbound

This paper cites Deep learning of representations: Looking forward,.

Optimizing Speech Multi-View Feature Fusion through Conditional Computation Deep learning of representations: Looking forward,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:35:13.544333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T20:35:13.035113Z digest=sha256:faa3f098d688fe5aa3d97a1140d49250a9dd02581bde72d7fbbc36d7e3ce255d

Observation d1f92b65-ee65-4a44-b930-b9bf5021c6f9 · outbound

This paper cites Depth-Adaptive Transformer.

Optimizing Speech Multi-View Feature Fusion through Conditional Computation Depth-Adaptive Transformer

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:13.040015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:13.040015Z digest=sha256:5eaac052374cd3b43b72788ffbb3055cbb232cd1c69d250c7f00ef96f37d493f

Observation 5f92795d-8300-4ff3-a32a-b55de0f7c3a0 · outbound

This paper cites LayerSkip: Enabling Early Exit Inference and Self-Speculative Decoding.

Optimizing Speech Multi-View Feature Fusion through Conditional Computation LayerSkip: Enabling Early Exit Inference and Self-Speculative Decoding

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:13.048303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:13.048303Z digest=sha256:afb79d9df6fca1badd4dbdcfffd84646df75c1abccabaa6af185909b8f48d2e2

Observation 1de2a624-7275-406c-99f7-68507eac3d18 · outbound

This paper cites Modeling task relationships in multi-task learning with multi- gate mixture-of-experts,.

Optimizing Speech Multi-View Feature Fusion through Conditional Computation Modeling task relationships in multi-task learning with multi- gate mixture-of-experts,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:35:13.525266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T20:35:13.054186Z digest=sha256:fc7943d861b150816b4ca84fb8bbbd98fcfe444956a568749a9267829bfabe15

Observation 263ec31d-b428-4805-8b60-79b344d77f15 · outbound

This paper cites Sparsely Activated Mixture-of-Experts are Robust Multi-Task Learners.

Optimizing Speech Multi-View Feature Fusion through Conditional Computation Sparsely Activated Mixture-of-Experts are Robust Multi-Task Learners

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:13.059024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:13.059024Z digest=sha256:00027405dd42ec2854dc5b6ec234583a5c55e9ffe028081f83bd8b6dfc863aa6

Observation fa28073f-17e4-43b5-b200-a7138ab9261d · outbound

This paper cites Stacked Acoustic-and-Textual Encoding: Integrating the Pre-trained Models into Speech Translation Encoders.

Optimizing Speech Multi-View Feature Fusion through Conditional Computation Stacked Acoustic-and-Textual Encoding: Integrating the Pre-trained Models into Speech Translation Encoders

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:13.064469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:13.064469Z digest=sha256:668eab84faaddd62dfc6226f7cdd5093eb6d65242c0e957c95e6e33f4fe5c335

Observation 83c81a68-fcf4-4208-8a2d-4c9dba5f3425 · outbound

This paper cites Soft Alignment of Modality Space for End-to- End Speech Translation,.

Optimizing Speech Multi-View Feature Fusion through Conditional Computation Soft Alignment of Modality Space for End-to- End Speech Translation,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:35:13.505664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T20:35:13.069007Z digest=sha256:578b2f0129afe0ebdc01d89670ee84f3b7d0b40d8324c6fc5f9c00e3072e2d82

Observation 0c0b2c8b-78c2-4d45-8c40-a2ae4576bdef · outbound

This paper cites Deep residual learning for image recognition,.

Optimizing Speech Multi-View Feature Fusion through Conditional Computation Deep residual learning for image recognition,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:13.073686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:13.073686Z digest=sha256:54b29abbcdde83e150ce1574e778d4ef39c9976e8d82b0e22e45e56f069c35d2

Observation c91f3c40-8959-479e-8e60-55284b580e68 · outbound

This paper cites Progres- sive layered extraction (ple): A novel multi-task learning (mtl) model for personalized recommendations,.

Optimizing Speech Multi-View Feature Fusion through Conditional Computation Progres- sive layered extraction (ple): A novel multi-task learning (mtl) model for personalized recommendations,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:35:13.472636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T20:35:13.078447Z digest=sha256:ca31acf98268a38f09c83977813c525e220e8907e74bdfe5b772ca26e1494072

Observation ae696b85-5208-455c-aa0f-73fe035f49be · outbound

This paper cites Gradient surgery for multi-task learning,.

Optimizing Speech Multi-View Feature Fusion through Conditional Computation Gradient surgery for multi-task learning,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:35:13.459433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T20:35:13.085525Z digest=sha256:9b7117bceca9e45539ea3527c677c564bcc870fc00e84b5fad268b9926704c73

Observation 713c1147-2693-4839-b9f0-74d40a04d96d · outbound

This paper cites Must-c: a multilingual speech translation corpus,.

Optimizing Speech Multi-View Feature Fusion through Conditional Computation Must-c: a multilingual speech translation corpus,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:35:13.445597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T20:35:13.095399Z digest=sha256:a02b259c7d930793546d3e9a826ebf6fc63af3a9e7737b3b1156bacd27b9839d

Observation 71a5f744-1695-4b27-8df8-d5357121a6c1 · outbound

This paper cites Librispeech: an asr corpus based on public domain audio books,.

Optimizing Speech Multi-View Feature Fusion through Conditional Computation Librispeech: an asr corpus based on public domain audio books,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:35:13.427050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T20:35:13.101348Z digest=sha256:6f6e1f308f97c1809594af7714fa0ee26f8eb0525432e6c69e526835045f4a6d

Observation 59eb1700-894c-4115-9694-734437d47877 · outbound

This paper cites Modality Adaption or Regularization? A Case Study on End-to-End Speech Translation.

Optimizing Speech Multi-View Feature Fusion through Conditional Computation Modality Adaption or Regularization? A Case Study on End-to-End Speech Translation

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-08-10T20:35:13.209587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T20:35:13.106571Z digest=sha256:d9f9d14fc783993cbd509e24e984cc202d0239239c7724b1d84d81b304d29a41

Observation 0b0fa77c-ad3e-4681-a185-8a1832ed8f8f · outbound

This paper cites Attention Is All You Need.

Optimizing Speech Multi-View Feature Fusion through Conditional Computation Attention Is All You Need

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:13.113877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:13.113877Z digest=sha256:8d1a5a0599d5ff10ad1e24c06295124038643723329b625009296b18f1850d83

Observation 1d095de8-3b52-489a-81f0-adfc5bb1981d · outbound

This paper cites A Call for Clarity in Reporting BLEU Scores.

Optimizing Speech Multi-View Feature Fusion through Conditional Computation A Call for Clarity in Reporting BLEU Scores

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:13.120441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:13.120441Z digest=sha256:17bb7de37169904bc4be09452e0a701e94eda58f3734aa95eb31c50aee69261e

Pith citing papers

No inbound Pith citation observations are available.