Pith. sign in

Paper Citation Record · LEDGER

Speech Emotion Detection Based on MFCC and CNN-LSTM Architecture

As of 21 August 2026, this Paper Citation Record lists 14 of 14 outbound references and 2 inbound Pith citation observations for arXiv:2501.10666.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.10666 v1

Coverage vector

measured 14 of 14 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T19:04:41.053883Z

measured 16 of 16 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T19:55:18.061325Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T00:35:10.333532Z

Reference resolution

14 of 14 outbound references displayed

  • verified exact0
  • verified fuzzy10
  • unresolved4
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4144881c-a92a-4420-b33b-2dda3b467629 · outbound

This paper cites The emotion detection services could enable users to have enhanced experiences to manipulate and adjust their real-time emotions, receiving feedback timely.

Speech Emotion Detection Based on MFCC and CNN-LSTM Architecture The emotion detection services could enable users to have enhanced experiences to manipulate and adjust their real-time emotions, receiving feedback timely

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:04:41.319161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T19:04:40.979061Z digest=sha256:73cd2f0c15b858db30f6ba224fb5bc56960b8d853019c5e4eb2b2351a5e64931

Observation ac050e0f-d042-4c50-aae7-6305e4a8f608 · outbound

This paper cites Dataset description and preprocessing For the dataset, two relatively common-used datasets are taken into consideration as the original speech signal.

Speech Emotion Detection Based on MFCC and CNN-LSTM Architecture Dataset description and preprocessing For the dataset, two relatively common-used datasets are taken into consideration as the original speech signal

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:04:41.302489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T19:04:40.985979Z digest=sha256:00c2803bb0df5391fd00a7e44b164fa1d128983d913637f9a1e8f665439094df

Observation 64041510-569d-4326-9538-c27f1d87d443 · outbound

This paper cites The model loss for training and test set Firstly, the basic training effects of the model using the architectures designed above are illustrated in Figure 4.

Speech Emotion Detection Based on MFCC and CNN-LSTM Architecture The model loss for training and test set Firstly, the basic training effects of the model using the architectures designed above are illustrated in Figure 4

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:04:41.284912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T19:04:40.991971Z digest=sha256:855fa03cbbf63f31e1102a7ee57e4dfc246125a9187eb854c3004633e36c92af

Observation ce8db745-848a-47c9-8857-378e19902917 · outbound

This paper cites Regarding the feature extraction, features from four aspects are adopted for further analysis and the MFCC processing provided by the Librosa is specifically vital.

Speech Emotion Detection Based on MFCC and CNN-LSTM Architecture Regarding the feature extraction, features from four aspects are adopted for further analysis and the MFCC processing provided by the Librosa is specifically vital

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:04:41.268423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T19:04:40.997685Z digest=sha256:b42b21e5ef4f19cdb99a491ad53f0f0314efdef121d84b5a72a14f561de9e04f

Observation fbac4517-f616-4e4a-aa39-5e25e249d1cb · outbound

This paper cites an unresolved cited work.

Speech Emotion Detection Based on MFCC and CNN-LSTM Architecture Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-10T19:04:41.251130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T19:04:41.003012Z digest=sha256:091b27c805900065c29682f47a5b6fdec4779257805c830b019346f3011feeee

Observation f17e1887-6c35-46d8-b31a-bca48a40c738 · outbound

This paper cites 2015 Is virtual reality emotionally arousing? Investigating five emotion inducing virtual park scenarios International journal of human-computer studies 82 pp 48-56.

Speech Emotion Detection Based on MFCC and CNN-LSTM Architecture 2015 Is virtual reality emotionally arousing? Investigating five emotion inducing virtual park scenarios International journal of human-computer studies 82 pp 48-56

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:04:41.234302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T19:04:41.008842Z digest=sha256:4425d8f0568d263556173113fa4a9a78fe323059435cdc924f9e6a2b961413f8

Observation b5209c21-5024-453f-aacb-153467950917 · outbound

This paper cites an unresolved cited work.

Speech Emotion Detection Based on MFCC and CNN-LSTM Architecture Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-10T19:04:41.216861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T19:04:41.014325Z digest=sha256:347b622e03bbfe12fc369ae314e9ae1db56bf99770524459b77a65322f8b089b

Observation c2b51610-eee0-4f7a-be0e-d068c6f54496 · outbound

This paper cites 2015 Performance evaluation of different support vector machine kernels for face emotion recognition 2015 SAI Intelligent Systems Conference (IntelliSys) pp 804-806.

Speech Emotion Detection Based on MFCC and CNN-LSTM Architecture 2015 Performance evaluation of different support vector machine kernels for face emotion recognition 2015 SAI Intelligent Systems Conference (IntelliSys) pp 804-806

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:04:41.200231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T19:04:41.019480Z digest=sha256:03f5a1785a689125621c3fe25cecefef07959fbbb1143243dacc82c933e46185

Observation c60ab30b-18d1-4fdb-b357-f014d436cf1d · outbound

This paper cites 2015 Improved emotion recognition using GMM-UBMs 2015 International Conference on Signal Processing and Communication Engineering Systems pp 53-57.

Speech Emotion Detection Based on MFCC and CNN-LSTM Architecture 2015 Improved emotion recognition using GMM-UBMs 2015 International Conference on Signal Processing and Communication Engineering Systems pp 53-57

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:04:41.181259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T19:04:41.025457Z digest=sha256:55951d7bd021cf977963b61fef35a6a29415168d5f236e26ee9b9a9d413b44f4

Observation bef11215-8d8e-42c6-9b99-a328543a7165 · outbound

This paper cites an unresolved cited work.

Speech Emotion Detection Based on MFCC and CNN-LSTM Architecture Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-10T19:04:41.161639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T19:04:41.031083Z digest=sha256:261ce14c7d042d653ff4b6ba9720959d418344ae363865b034d8f0a5eaa53d6e

Observation 51ce6d2d-c645-462f-8d79-95d42c766d33 · outbound

This paper cites 2015 librosa: Audio and music signal analysis in python Proceedings of the 14th python in science conference Vol 8 pp 18-25.

Speech Emotion Detection Based on MFCC and CNN-LSTM Architecture 2015 librosa: Audio and music signal analysis in python Proceedings of the 14th python in science conference Vol 8 pp 18-25

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:04:41.145025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T19:04:41.036154Z digest=sha256:7514e49397f7c8d95ba0bc589a1c60e8fc5ee0a5b6099b2c9a1f2897c33150ca

Observation 1c9e2c5b-228f-47b0-bb5a-e43d1d46fecc · outbound

This paper cites an unresolved cited work.

Speech Emotion Detection Based on MFCC and CNN-LSTM Architecture Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-10T19:04:41.128465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T19:04:41.041789Z digest=sha256:7b3f5e88427c5bf70a9bc262500fcbbf582dd75f0ee84ec5317f34e765b1952b

Observation 2f3a092d-0e79-4250-a45c-87f9b6bcbee2 · outbound

This paper cites 2015 Long-term recurrent convolutional networks for visual recognition and description Proceedings of the IEEE conference on computer vision and pattern recognition pp 2625- 2634.

Speech Emotion Detection Based on MFCC and CNN-LSTM Architecture 2015 Long-term recurrent convolutional networks for visual recognition and description Proceedings of the IEEE conference on computer vision and pattern recognition pp 2625- 2634

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:04:41.111988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T19:04:41.047194Z digest=sha256:e3406d7aee5b0b2189cddbf54f75712e5b268f03e1f77ed4e35f3fefb20974ad

Observation 4abf6c28-5db3-4931-9364-7b7d3232e262 · outbound

This paper cites 2015 Classification of emotions of angry and disgust SmartCR 5(3) pp 151-158.

Speech Emotion Detection Based on MFCC and CNN-LSTM Architecture 2015 Classification of emotions of angry and disgust SmartCR 5(3) pp 151-158

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:04:41.094805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T19:04:41.053883Z digest=sha256:650426e58b6611e384888324d073f61efd5a59524a383affdf989896e79132e4

Pith citing papers

Observation 9274084a-bcfe-402f-8827-62114a6f2b12 · inbound

Low Resource Multimodal Translation of Nepali Spoken Words into Emotion-Conditioned Sign Language Avatars cites this paper.

Low Resource Multimodal Translation of Nepali Spoken Words into Emotion-Conditioned Sign Language Avatars Speech Emotion Detection Based on MFCC and CNN-LSTM Architecture

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-07-01T00:35:10.334848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-07-01T00:29:09.303984Z digest=sha256:9a65227799f9c08adc153871dc28d49314b8840221918edd2bffb92ce66768de

Observation 9e87df13-b8c2-4bf1-a5e3-5575ff89b164 · inbound

Explainable Lightweight Compact Deep Models for Speech Emotion Recognition cites this paper.

Explainable Lightweight Compact Deep Models for Speech Emotion Recognition Speech Emotion Detection Based on MFCC and CNN-LSTM Architecture

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-01T19:55:18.061325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T19:55:18.061325Z digest=sha256:7a631006a629f970c954697acf701a45aef9738c2c20df59d2673d08cc3bfce3