Pith. sign in

Paper Citation Record ยท LEDGER

Towards Efficient and Multifaceted Computer-assisted Pronunciation Training Leveraging Hierarchical Selective State Space Model and Decoupled Cross-entropy Loss

As of 9 August 2026, this Paper Citation Record lists 18 of 18 outbound references and 1 inbound Pith citation observation for arXiv:2502.07575.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.07575 v2

Coverage vector

measured 18 of 18 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T12:19:08.006769Z

measured 19 of 19 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T23:10:57.314825Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-06T23:10:57.510789Z

Reference resolution

18 of 18 outbound references displayed

  • verified exact0
  • verified fuzzy8
  • unresolved10
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f2126658-2576-4a79-b045-59896f80d958 ยท outbound

This paper cites However, it is generally expected that a full-fledged CAPT system should perform both functionalities simultaneously and efficiently.

Towards Efficient and Multifaceted Computer-assisted Pronunciation Training Leveraging Hierarchical Selective State Space Model and Decoupled Cross-entropy Loss However, it is generally expected that a full-fledged CAPT system should perform both functionalities simultaneously and efficiently

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:19:08.392519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T12:19:07.921568Z digest=sha256:84de8deac3fee1cbeacf4c6b4a357d053dcfaf3df250f68e7905e0604ac44ccf

Observation 5f041df5-8e2a-46a8-b284-765efafd6e7c ยท outbound

This paper cites These modules collectively generate the corresponding aspect score sequence ๐ฌ๐‘” for each linguistic granularity ๐‘”, as well as the phonetic error states ๐ž and diagnosis ๐ฒ.

Towards Efficient and Multifaceted Computer-assisted Pronunciation Training Leveraging Hierarchical Selective State Space Model and Decoupled Cross-entropy Loss These modules collectively generate the corresponding aspect score sequence ๐ฌ๐‘” for each linguistic granularity ๐‘”, as well as the phonetic error states ๐ž and diagnosis ๐ฒ

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:19:08.347525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T12:19:07.937362Z digest=sha256:336d32eb0991415b9ce564301ae10a390d4ce36a50fdd04e775a6c592efe8d3e

Observation d43c3eba-9f01-4768-8695-0434ed7c6a80 ยท outbound

This paper cites These errors usually have clear-cut distinctions between correct and incorrect ones, and can be easily quantified through deletions, substitutions, and insertions.

Towards Efficient and Multifaceted Computer-assisted Pronunciation Training Leveraging Hierarchical Selective State Space Model and Decoupled Cross-entropy Loss These errors usually have clear-cut distinctions between correct and incorrect ones, and can be easily quantified through deletions, substitutions, and insertions

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:19:08.362769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T12:19:07.932079Z digest=sha256:e99d829095e2826fd1981c708f2b7bd3eaf0cc09c18e48f1c31d6b052722839f

Observation d44e73af-d10a-4229-8d33-c02c717ac438 ยท outbound

This paper cites Notably, there are several studies investigating the bidirectional processing of Mamba (Liang et al., 2024; Zhang et al., 2024; Jiang et al., 2024).

Towards Efficient and Multifaceted Computer-assisted Pronunciation Training Leveraging Hierarchical Selective State Space Model and Decoupled Cross-entropy Loss Notably, there are several studies investigating the bidirectional processing of Mamba (Liang et al., 2024; Zhang et al., 2024; Jiang et al., 2024)

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:19:08.332810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T12:19:07.942760Z digest=sha256:14d7a096807d9e725ebcecbe57c2416252641bc8a6229ad7da22af12ae099106

Observation 7a149287-4e5f-4743-8832-f9df68db3bc1 ยท outbound

This paper cites The APA module contains one regressor that aims to predict the phone-level aspect score ๐‘ 0๐‘”๐‘โ„Ž๐‘›(accuracy).

Towards Efficient and Multifaceted Computer-assisted Pronunciation Training Leveraging Hierarchical Selective State Space Model and Decoupled Cross-entropy Loss The APA module contains one regressor that aims to predict the phone-level aspect score ๐‘ 0๐‘”๐‘โ„Ž๐‘›(accuracy)

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:19:08.303935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T12:19:07.953380Z digest=sha256:ab507efe515e0ed983eb8fd0a445256ebb7d367c70512046733865d1aa402396

Observation d33575be-4913-41ca-b1c2-f55852cc5c33 ยท outbound

This paper cites Mamba: Linear-Time Sequence Modeling with Selective State Spaces.

Towards Efficient and Multifaceted Computer-assisted Pronunciation Training Leveraging Hierarchical Selective State Space Model and Decoupled Cross-entropy Loss Mamba: Linear-Time Sequence Modeling with Selective State Spaces

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-08T12:19:07.976059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:19:07.976059Z digest=sha256:c4f0f6c6cf156987bfcd537e91527bc990cf52dd1e3e9907e9853b04a7dc8f83

Observation d03425cd-6d26-4c28-b4e7-88107123665a ยท outbound

This paper cites Bi-Mamba+: Bidirectional Mamba for Time Series Forecasting.

Towards Efficient and Multifaceted Computer-assisted Pronunciation Training Leveraging Hierarchical Selective State Space Model and Decoupled Cross-entropy Loss Bi-Mamba+: Bidirectional Mamba for Time Series Forecasting

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-08T12:19:07.985835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:19:07.985835Z digest=sha256:04a8736fcc3713bbbdeef41db45a29ef9cd5038551f6d1f85cc2461b93ee6900

Observation 0d299bed-cb26-457e-b369-8990e051a3ef ยท outbound

This paper cites Vision Mamba: Efficient Visual Representation Learning with Bidirectional State Space Model.

Towards Efficient and Multifaceted Computer-assisted Pronunciation Training Leveraging Hierarchical Selective State Space Model and Decoupled Cross-entropy Loss Vision Mamba: Efficient Visual Representation Learning with Bidirectional State Space Model

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-08T12:19:07.990806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:19:07.990806Z digest=sha256:5816137297431a9841e874e32c1d04b98bc540da9db26674c5de2b9c1330461f

Observation c6c79692-9b67-4d9b-80da-5f9a76193c3f ยท outbound

This paper cites Mamba in Speech: Towards an Alternative to Self-Attention.

Towards Efficient and Multifaceted Computer-assisted Pronunciation Training Leveraging Hierarchical Selective State Space Model and Decoupled Cross-entropy Loss Mamba in Speech: Towards an Alternative to Self-Attention

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-08T12:19:07.995703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:19:07.995703Z digest=sha256:a1ba8760fbd7bf15323676496ecbdc36f79423f1ea8b0d3a9825753c868f310e

Observation fc9f1182-ca35-4b55-82ad-d3eb0038fcfd ยท outbound

This paper cites The combining weights ๐œ”๐‘” for APA loss are uniformly set to 1.0 for each granularity level ๐‘”.

Towards Efficient and Multifaceted Computer-assisted Pronunciation Training Leveraging Hierarchical Selective State Space Model and Decoupled Cross-entropy Loss The combining weights ๐œ”๐‘” for APA loss are uniformly set to 1.0 for each granularity level ๐‘”

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:19:08.241622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T12:19:08.001187Z digest=sha256:68a2654a8fc8e0c9a45e371625137e20e8ff728316d46e3a7be17e9606f8b1de

Observation ad280e6f-533a-4f2b-b9a7-04ed58894fec ยท outbound

This paper cites an unresolved cited work.

Towards Efficient and Multifaceted Computer-assisted Pronunciation Training Leveraging Hierarchical Selective State Space Model and Decoupled Cross-entropy Loss Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-08T12:19:08.226328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T12:19:08.006769Z digest=sha256:e3f9ca458ccad3bf1349ebab86141ec9dce662b1a02bde28d82d094111ceb4cc

Observation 0113c96c-b848-4660-8370-04299622ba88 ยท outbound

This paper cites reading-aloud.

Towards Efficient and Multifaceted Computer-assisted Pronunciation Training Leveraging Hierarchical Selective State Space Model and Decoupled Cross-entropy Loss reading-aloud

Reference 2009

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:19:08.377896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T12:19:07.927281Z digest=sha256:197ad6498ea4b74af3caf40f8acd66ecc0fdd6cde0c6f0a1d8022648a083a373

Observation 79022266-41fe-4f0e-8d26-41c3171136ea ยท outbound

This paper cites ETS Research Report Series 2015(1):1โ€“11.

Towards Efficient and Multifaceted Computer-assisted Pronunciation Training Leveraging Hierarchical Selective State Space Model and Decoupled Cross-entropy Loss ETS Research Report Series 2015(1):1โ€“11

Reference 2015

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:19:08.256312Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T12:19:07.967573Z digest=sha256:8a4115363420dd10cd9ef7a41e04f52814fc577c7f7844cd9cbe253a8f6bf9ae

Observation e3ff5edd-324f-47b2-b4ca-24e708ea7001 ยท outbound

This paper cites an unresolved cited work.

Towards Efficient and Multifaceted Computer-assisted Pronunciation Training Leveraging Hierarchical Selective State Space Model and Decoupled Cross-entropy Loss Unresolved cited work

Reference 2017

Resolution
unresolved
raw_fallback, observed 2026-08-08T12:19:08.287512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T12:19:07.958489Z digest=sha256:d0eb56c66c117e596d8c1a074bae420352d80bdbbbaaa524a089ca03ba3e0700

Observation a861861f-8133-43c8-b8a7-bc5772cc47c0 ยท outbound

This paper cites A Full Text-Dependent End to End Mispronunciation Detection and Diagnosis with Easy Data Augmentation Techniques.

Towards Efficient and Multifaceted Computer-assisted Pronunciation Training Leveraging Hierarchical Selective State Space Model and Decoupled Cross-entropy Loss A Full Text-Dependent End to End Mispronunciation Detection and Diagnosis with Easy Data Augmentation Techniques

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-08T12:19:07.971301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:19:07.971301Z digest=sha256:a573d08114118f847d79f38ec335c00cfe508022f6e23b85d70a9dd7d8715cf2

Observation ef42ce8a-df5a-4f20-9090-d87136f429f4 ยท outbound

This paper cites an unresolved cited work.

Towards Efficient and Multifaceted Computer-assisted Pronunciation Training Leveraging Hierarchical Selective State Space Model and Decoupled Cross-entropy Loss Unresolved cited work

Reference 2022

Resolution
unresolved
raw_fallback, observed 2026-08-08T12:19:08.272001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T12:19:07.963458Z digest=sha256:ffcab167564c05d2dc56f1dad22f45d97f6413abc0b7c4b45e4e30803e3e704b

Observation 399c6f3b-d3df-4dd3-b4a5-4f9650e4de89 ยท outbound

This paper cites an unresolved cited work.

Towards Efficient and Multifaceted Computer-assisted Pronunciation Training Leveraging Hierarchical Selective State Space Model and Decoupled Cross-entropy Loss Unresolved cited work

Reference 2023

Resolution
unresolved
raw_fallback, observed 2026-08-08T12:19:08.319060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T12:19:07.948120Z digest=sha256:bb67a6b22b8396d627a3fabf44a068c460d36422477783e1cd60478ad025f321

Observation f86411e5-577b-4b4b-ad97-12f63edb952f ยท outbound

This paper cites Dual-path Mamba: Short and Long-term Bidirectional Selective Structured State Space Models for Speech Separation.

Towards Efficient and Multifaceted Computer-assisted Pronunciation Training Leveraging Hierarchical Selective State Space Model and Decoupled Cross-entropy Loss Dual-path Mamba: Short and Long-term Bidirectional Selective Structured State Space Models for Speech Separation

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-08T12:19:07.980987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:19:07.980987Z digest=sha256:58d8cbedeeead4ce5ab124ee8ebc82565b0f6b89d54fb8634d003d9f4dc07076

Pith citing papers

Observation a2f6d2ec-6c8c-4498-9198-851eefda73ac ยท inbound

JCAPT: A Joint Modeling Approach for CAPT cites this paper.

JCAPT: A Joint Modeling Approach for CAPT Towards Efficient and Multifaceted Computer-assisted Pronunciation Training Leveraging Hierarchical Selective State Space Model and Decoupled Cross-entropy Loss

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-08-06T23:10:57.516689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T23:10:57.314825Z digest=sha256:ce7ce3dad43dba439619f8c22a5557df82c51b3491c1c7aee2d297395c2cc47d